Aivora
arXivAI ResearchAdvanced

Distilling Graph Geometry: Bridging the GNN-to-MLP Knowledge Gap

蒸餾圖幾何學:縮補 GNN 到 MLP 知識蒸餾的幾何黑洞

2 min read
Distilling Graph Geometry: Bridging the GNN-to-MLP Knowledge Gap
The 30-second version

While GNNs are powerful, their dependency on graph structures during inference causes latency. Distilling GNNs into graph-free MLPs is a common solution, but existing methods suffer from spectral underfit in sparse graphs and spectral overfit in dense graphs. To address this, G²MLP leverages Ollivier-Ricci curvature to pinpoint where geometric errors occur, dynamically balancing prediction-level and representation-level distillation. The resulting MLP requires zero graph access at inference while retaining GNN-level performance.

Key points

01

Diagnosing Spectral Failure Modes

Identifies that student MLPs suffer from spectral underfit on sparse graphs and spectral overfit on dense graphs, losing key geometric structures.

02

Geometry-Aware Distillation (G²MLP)

Introduces an energy-weighted alignment guided by Ollivier-Ricci curvature to accurately preserve graph geometry in the student's representation space.

03

Graph-Free Inference Deployment

The deployed model is a standard MLP requiring zero graph access or message passing during inference.

04

Broad Versatility

Seamlessly scales across node classification, link prediction, and complex teacher models like Graph Transformers.

How it works

G²MLP Geometry-Aware Distillation Flow
Forward passAnalyze structureProvide predictions & representationsAllocate supervision weightsGeometry-guided distillationDeploy graph-free modelGraph & Node FeaturesGNN Teacher ModelOllivier-RicciCurvatureDynamic SupervisionAlignmentMLP Student (Training)Standard MLP(Graph-free Inference)

Why it matters

This work overcomes the high latency of GNNs in production environments caused by graph dependency. Through G²MLP, developers can compress complex GNNs into highly efficient MLPs. It bridges the gap between graph geometry and Euclidean space, providing a practical and theoretically grounded deployment solution for low-latency graph learning applications such as real-time recommendation and large-scale social network analysis.

Who it affects

  • AI Developer
  • AI Researcher

How to use it

  1. 1Real-time Recommendation Systems: Transferring graph-based knowledge to ultra-low latency MLPs for instant product or friend recommendations.
  2. 2Large-scale Link Prediction: Performing graph-free relationship predictions in edge devices or high-throughput environments.

Limitations & caveats

  • Computing Ollivier-Ricci curvature during training can introduce additional computational overhead on massive graph datasets.
  • Dynamic graph structures require retraining or re-distillation to maintain accurate representation of changing geometries.

Related