New LoRA Skills Should Read but Never Write: READ Solves Adapter Fusion Interference
新增 LoRA 技能唯讀不寫:READ 解決多配接器融合干擾
Merging multiple independently trained LoRA adapters often degrades performance due to parameter interference. This paper identifies two hidden culprits: arbitrary factorization coordinates and bidirectional coupling that overwrites old skills. To solve this, the authors propose READ (Read-only Expansion of Adapter Deltas). READ transforms each adapter into a balanced canonical form and enforces a strict one-way coupling: new skills can only read old skills' input subspaces but never write to their outputs. READ outperforms strong baselines by over 20 points on SuperGLUE and 7 points on domain benchmarks.
Key points
Identifying Interference Root Causes
The study reveals that interference in merged LoRAs stems from arbitrary factorization coordinates and bidirectional coupling that disrupts pre-existing skills.
Balanced Canonical Form
READ rewrites each adapter into a balanced canonical form, ensuring consistent coordinate baselines for interactions between different adapters.
One-Way Read-Only Coupling
New skills can only read the input subspaces of old skills but are strictly blocked from writing to their output subspaces, avoiding degradation.
Zero Inference Overhead
The composed updates can be folded directly back into the base model weights, eliminating the need for runtime routing or extra computation.
How it works
| 傳統 LoRA 合併 (Traditional) | READ (本研究方法) | |
|---|---|---|
| Coupling Direction | 雙向耦合(容易互相覆寫與干擾) | 單向唯讀(新技能唯讀舊技能,不干擾舊輸出) |
| Factorization | 任意、不一致的座標系統 | 標準化平衡型式 (Balanced Canonical Form) |
| Inference Cost | 高(若使用路由/MoE)或 表現不佳(直接相加) | 零成本(可直接摺疊融入基礎模型權重中) |
| Skill Retention | 差(技能易在合併過程中流失) | 極佳(舊技能計算路徑完全不受影響) |
Why it matters
Traditionally, combining multiple specialized skills in an LLM required expensive multi-task retraining or complex MoE-style routing that added inference latency. READ offers an elegant mathematical solution to "stack" or "hot-plug" new skills at near-zero cost without damaging existing capabilities. This is highly valuable for enterprise LLM deployments requiring continuous learning and modular feature expansion.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Modular AI Skill Stacking
- 2Lifelong Learning without Catastrophic Forgetting
Limitations & caveats
- Requires access to and transformation of the original LoRA adapter weights
- Primarily evaluated on sequential skill addition; behavior with extremely high numbers of concurrent adapters is untested
Related
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
自我監督信心訓練:免於刻意「學會停止」即可提升 LLM 推理效率
Researchers found that training reasoning models to predict their own confidence at intermediate steps naturally reduces generated tokens by up to 25% at matched accuracy, without explicitly optimizing for length or stopping.
First-Order Stationarity of Reverse Diffusions: Bridging Optimization and Sampling
逆向擴散的一階駐點性:連結優化與生成採樣的數學機制
This research establishes a first-order optimization theory for diffusion models, proving that SDE-based reverse Langevin diffusions contract Fisher divergences exponentially under strongly convex noising.
Statistical Attribute Alignment for Black-Box Generative AI via Output Post-Processing
黑盒生成式 AI 的統計屬性對齊:透過輸出後處理實現公平與多樣性
This paper introduces post-processing algorithms to align the attribute distribution of black-box generative AI outputs with user-specified targets using a mathematically minimized number of queries.