Distilling Graph Geometry: Bridging the GNN-to-MLP Knowledge Gap
蒸餾圖幾何學:縮補 GNN 到 MLP 知識蒸餾的幾何黑洞
While GNNs are powerful, their dependency on graph structures during inference causes latency. Distilling GNNs into graph-free MLPs is a common solution, but existing methods suffer from spectral underfit in sparse graphs and spectral overfit in dense graphs. To address this, G²MLP leverages Ollivier-Ricci curvature to pinpoint where geometric errors occur, dynamically balancing prediction-level and representation-level distillation. The resulting MLP requires zero graph access at inference while retaining GNN-level performance.
Key points
Diagnosing Spectral Failure Modes
Identifies that student MLPs suffer from spectral underfit on sparse graphs and spectral overfit on dense graphs, losing key geometric structures.
Geometry-Aware Distillation (G²MLP)
Introduces an energy-weighted alignment guided by Ollivier-Ricci curvature to accurately preserve graph geometry in the student's representation space.
Graph-Free Inference Deployment
The deployed model is a standard MLP requiring zero graph access or message passing during inference.
Broad Versatility
Seamlessly scales across node classification, link prediction, and complex teacher models like Graph Transformers.
How it works
Why it matters
This work overcomes the high latency of GNNs in production environments caused by graph dependency. Through G²MLP, developers can compress complex GNNs into highly efficient MLPs. It bridges the gap between graph geometry and Euclidean space, providing a practical and theoretically grounded deployment solution for low-latency graph learning applications such as real-time recommendation and large-scale social network analysis.
Who it affects
- AI Developer
- AI Researcher
How to use it
- 1Real-time Recommendation Systems: Transferring graph-based knowledge to ultra-low latency MLPs for instant product or friend recommendations.
- 2Large-scale Link Prediction: Performing graph-free relationship predictions in edge devices or high-throughput environments.
Limitations & caveats
- Computing Ollivier-Ricci curvature during training can introduce additional computational overhead on massive graph datasets.
- Dynamic graph structures require retraining or re-distillation to maintain accurate representation of changing geometries.
Related
Building Persistent 3D Object Memory: How Ledger Tracks Objects from Egocentric Videos
打造過目不忘的 3D 空間記憶:Ledger 如何透過第一人稱影片追蹤隱形物體
Researchers introduce Ledger, a framework that builds a persistent 3D object memory from egocentric videos, significantly improving spatial question-answering accuracy for embodied agents.
Decoupling Exploration from Optimization: How ExpDis Boosts LLM Reasoning and Solution Diversity
探索與優化解耦:全新強化學習框架 ExpDis 提升大語言模型的推理多元性
The ExpDis framework decouples exploration from optimization in RLVR. By training explorers with novelty bonuses and distilling filtered trajectories into a student model, it prevents model degradation while fostering diverse reasoning.
Clipped Decentralized SGD: Achieving Optimal Convergence and Linear Speed-Up Under Heavy-Tailed Noise
去中心化 SGD 克服重尾雜訊:梯度裁剪如何實現最佳收斂與線性加速
This study proves that clipped decentralized SGD (DSGD) achieves order-optimal convergence rates and linear speed-up under heavy-tailed noise for non-convex optimization.