Common-Mode Collapse and Recovery in Direct Feedback Alignment
直接回饋對齊(DFA)中的共模崩潰與恢復機制
Direct Feedback Alignment (DFA) trains networks using fixed random projections of error as an alternative to backpropagation. This study reveals that with tanh activations, DFA often stalls early due to 'common-mode collapse'—where shared error across inputs drives units to saturation. Using exact mean-covariance decomposition, the authors analyze this phenomenon on MNIST and CIFAR-10, demonstrating that Adam optimization, readout calibration, or batch mean subtraction can successfully prevent or recover from this collapse.
Key points
Origin of Common-Mode Collapse
The error's common mode shared across inputs creates a rank-one update that drives tanh units toward saturation and stalls training.
Optimizer & Calibration Impact
While Adam optimizer leads to deeper collapse, it learns faster than SGD. Calibrating the baseline readout to the class prior suppresses the collapse.
Preventing Collapse via Mean Subtraction
Subtracting the signal's batch mean prevents sustained collapse and improves the overall learning performance.
Why it matters
DFA is a biologically plausible and hardware-friendly alternative to backpropagation. This research diagnoses the mathematical limitations (common-mode collapse) of DFA during early training and provides concrete solutions, guiding the design of more stable and efficient non-backprop neuromorphic algorithms.
Who it affects
- AI Researcher
- AI Developer
How to use it
- 1Improving non-backpropagation neural network training algorithms on neuromorphic hardware.
- 2Optimizing low-power on-chip learning systems based on DFA in edge devices.
Limitations & caveats
- The severity and mitigation cost of the collapse highly depend on the readout setup, choice of optimizer, and input statistics.
- The study focuses primarily on tanh activations and specific image classification tasks; generalization to other architectures like Transformers remains to be fully explored.
Related
FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
FurE:免用動物毛髮資料集,實現 10 倍加速的 3D 動物毛髮重建技術
FurE is an efficient 3D animal fur reconstruction method that leverages a human-hair trained PCA decoder and Gaussian Frosting to achieve 10x faster, highly detailed, and editable groom reconstruction without animal datasets.
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
自我監督信心訓練:免於刻意「學會停止」即可提升 LLM 推理效率
Researchers found that training reasoning models to predict their own confidence at intermediate steps naturally reduces generated tokens by up to 25% at matched accuracy, without explicitly optimizing for length or stopping.
First-Order Stationarity of Reverse Diffusions: Bridging Optimization and Sampling
逆向擴散的一階駐點性:連結優化與生成採樣的數學機制
This research establishes a first-order optimization theory for diffusion models, proving that SDE-based reverse Langevin diffusions contract Fisher divergences exponentially under strongly convex noising.