Untying Embeddings in LLMs Improves Privacy-Preserving DP-SGD Fine-Tuning Utility and Memory
差分隱私訓練新發現:解除權重共享能為 LLM 提升 4.74% 準確度並降低 60% 記憶體
This study investigates if weight tying—sharing weights between input and output embeddings—remains beneficial for decoder-only LLMs trained under DP-SGD. Evaluating GPT2 and DistilGPT2 on SST-2, QNLI, and QQP, the researchers found that untied models consistently outperform tied ones, yielding up to 4.74% higher accuracy. Crucially, untying embeddings enables memory-efficient ghost clipping by removing shared-parameter interactions, resulting in over 60% reduction in training memory usage.
Key points
Significant Accuracy Boost
Untying embeddings consistently improves model utility under DP-SGD, yielding gains of up to 4.74 percentage points in accuracy.
Over 60% Memory Savings
Untied models achieve over 60% lower memory usage during DP-SGD fine-tuning by allowing standard, efficient ghost clipping.
Challenging Architecture Norms
Weight tying, originally designed for non-private parameter efficiency, is shown to be counterproductive in private training settings.
How it works
| 權重共享 (Weight Tying) | 解除權重共享 (Untied Embeddings) | |
|---|---|---|
| DP-SGD Accuracy | 較低 (Lower) | 最高提升 4.74% (Up to 4.74% Higher) |
| Memory Usage | 較高 (Higher) | 降低超過 60% (Over 60% Lower) |
| Ghost Clipping Support | 不相容/計算複雜且失效 (Negated due to parameter sharing) | 完美相容且高效 (Fully compatible and highly efficient) |
| Original Purpose | 非隱私下的參數效率 (Non-private parameter efficiency) | 隱私保護下的高實用與高效能 (Privacy-centric utility and efficiency) |
Why it matters
This research provides key architectural guidance for privacy-preserving LLM fine-tuning. While weight tying is standard for saving parameters in non-private training, it degrades performance and blocks memory-saving techniques like ghost clipping under DP-SGD. By simply untying embeddings, organizations can achieve both higher accuracy and over 60% lower memory footprint, making private LLM training much more practical and accessible on standard hardware.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Privacy-preserving LLM fine-tuning using DP-SGD on sensitive medical or financial data
- 2Memory-efficient private model training on consumer-grade GPUs using ghost clipping
Limitations & caveats
- Evaluations are primary conducted on GPT-2 and DistilGPT-2 architectures; validation on massive billion-scale models remains to be explored.
- The benchmarks are limited to classification and NLU tasks (SST-2, QNLI, QQP), leaving open generative tasks unassessed.
Related

Overcoming Generative Recommender Latency: Deploying HSTU Models with NVIDIA Dynamo-Triton and PyTorch AOTI
突破生成式推薦延遲瓶頸:NVIDIA Dynamo-Triton 與 PyTorch AOTI 部署 HSTU 模型實戰
Learn how to deploy HSTU generative recommenders using NVIDIA Dynamo-Triton, PyTorch AOTI, and FlexKV caching to achieve up to a 5.93x speedup on Blackwell GPUs.
Ranking-PE: Prompt Optimization for Multimodal Clinical Diagnosis under Extreme Class Imbalance
臨床診斷 MLLM 提示詞優化:Ranking-PE 解決醫療資料極端不平衡問題
This paper introduces Ranking-PE, a ranking-aware prompt optimization framework that shifts MLLM adaptation from accuracy-based to AUROC-based ranking, resolving class imbalance in clinical diagnostics.
Fixing the "Timing Shortcut": A Breakthrough in Non-Invasive Brain-to-Text Decoding
排除「時間捷徑」漏洞:非侵入式腦機介面解碼技術的新突破
Researchers revealed that recent breakthroughs in non-invasive brain-to-text decoding relied on a "timing shortcut" of word durations rather than actual brain signals. Their SimpleB2T method eliminates this shortcut, slashing the word error rate to 36.6%.