Weight Pair Encoding (WeightPE): Inducing a Smaller Grammar in Neural Network Weights
權重配對編碼(WeightPE):將類神經網路權重壓縮為更小、更具彈性的語法結構
Traditional network compression relies on flat, fixed-size codebooks. WeightPE takes a novel approach by flattening int8 weights into a string and placing a lossy Re-Pair compressor inside a straight-through estimator (STE). Under a global L2 budget, it aligns near-matching weight patterns into exact duplicates, training the network to adaptively induce a smaller grammar. Experiments on ViT models show that WeightPE reduces grammar size to 38%-43% of standard QAT, costing only 1.1 to 1.9 accuracy points while generalizing to non-targeted compressors.
Key points
Grammar Size as Objective
This represents the first time grammar size has been used as an explicit training objective for neural network weights.
Hierarchical & Variable-Length
Unlike flat, fixed-size codebooks, a grammar supports variable-length patterns and reuses them hierarchically inside larger structures.
Lossy Re-Pair via STE
Integrates a lossy Re-Pair compressor inside a straight-through estimator to force near-matching weight patterns to become identical within an L2 budget.
Cross-Compressor Generalization
The optimized weights generalize effectively to other grammar compressors (e.g., LZ78, SEQUITUR) that were not targeted during training.
How it works
Why it matters
This research provides a novel paradigm for model compression and deployment on resource-constrained edge devices. By optimizing network weights into highly compressible grammar structures, it drastically lowers storage requirements and over-the-air transmission bandwidth. Its cross-compressor generalization ensures wide compatibility with different hardware-level decompression algorithms.
Who it affects
- AI Researcher
- AI Developer
- Enterprise Leader
How to use it
- 1Edge Device Deployment: Drastically reducing the storage footprint of heavy models like Vision Transformers (ViTs) on embedded chips and mobile devices.
- 2Bandwidth-Efficient Updates: Utilizing grammar compression to minimize bandwidth consumption during over-the-air (OTA) model updates.
Limitations & caveats
- It introduces a minor performance trade-off, leading to a loss of 1.1 to 1.9 accuracy points on ViT models.
- The evaluation is currently limited to int8 quantized MLP weights on CIFAR-10, requiring further validation on large-scale language models (LLMs).
Related
FurE: 10x Faster 3D Animal Fur Reconstruction Without Animal Datasets
FurE:免用動物毛髮資料集,實現 10 倍加速的 3D 動物毛髮重建技術
FurE is an efficient 3D animal fur reconstruction method that leverages a human-hair trained PCA decoder and Gaussian Frosting to achieve 10x faster, highly detailed, and editable groom reconstruction without animal datasets.
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency
自我監督信心訓練:免於刻意「學會停止」即可提升 LLM 推理效率
Researchers found that training reasoning models to predict their own confidence at intermediate steps naturally reduces generated tokens by up to 25% at matched accuracy, without explicitly optimizing for length or stopping.
First-Order Stationarity of Reverse Diffusions: Bridging Optimization and Sampling
逆向擴散的一階駐點性:連結優化與生成採樣的數學機制
This research establishes a first-order optimization theory for diffusion models, proving that SDE-based reverse Langevin diffusions contract Fisher divergences exponentially under strongly convex noising.