The problem
Photonic neural networks compute with light: a mesh of Mach-Zehnder interferometers (MZIs) performs matrix multiplications at the speed of light and with very little energy. Training them is the hard part. Backpropagation needs a precise digital model of the chip, and real chips never match their model.
Training on the chip itself, using only forward measurements, avoids that mismatch. A recent method, FFzero, combines layer-local learning with zeroth-order gradient estimates and showed promising results, but only in an idealised simulation.
What I did
- Built a physically grounded model of an MZI mesh in Clements configuration with thermo-optic phase shifters, from the single MZI transfer matrix up to the full network.
- Added the non-idealities of real hardware: fabrication variability (Monte Carlo), insertion loss and coupler imbalance, detector noise in a balanced coherent receiver, and thermal crosstalk between heaters.
- Designed a controlled comparison: local (layer-wise) and global training with the same zeroth-order estimator, so that locality is the only variable.
- Anchored the simulator to a published single-chip photonic neural network result before running the sweeps.
What I found
[SUMMARY OF THE MAIN RESULTS, 2–3 SENTENCES]
Why it matters
[WHAT THIS MEANS FOR ON-CHIP TRAINING OF REAL PHOTONIC HARDWARE]