The Missing Primitive: Diagnosing and Repairing Structural Mathematical Reasoning in LLMs
LLM 數學推理的「失落基元」:如何診斷並修復模型缺乏的結構化理解?
While LLMs perform well on math problems, their structural understanding remains questionable. This paper introduces 'Mathematical Primitives' and a new benchmark to evaluate reasoning across four dimensions: Discovery, Generation, Digestion, and Execution. Diagnostics show that raw accuracy masks specific capability gaps, and 'Discovery' is the primary bottleneck. To address this, the authors develop a primitive-privileged self-distillation framework that successfully transfers structured reasoning capabilities to student models, consistently boosting performance across benchmarks.
Key points
Introducing Mathematical Primitives
Proposes the concept of Mathematical Primitives to systematically probe whether LLMs possess underlying structural understanding of mathematical problems.
Four-Dimensional Diagnostic Framework
Dissects math reasoning into Discovery, Generation, Digestion, and Execution, revealing that surface accuracy often hides internal capability gaps.
Discovery as the Main Bottleneck
Diagnostics reveal that 'Discovery' is the dominant bottleneck in LLM reasoning, yet discovery-limited failures are the most amenable to post-training repair.
Primitive-Guided Self-Distillation
Introduces a primitive-privileged self-distillation framework to selectively transfer guided reasoning capabilities to student models, boosting performance across scales.
Why it matters
This research exposes the gap between surface-level math accuracy and deep structural comprehension in LLMs. By dissecting reasoning into four distinct capabilities, developers can pinpoint exact model weaknesses. Crucially, proving that the dominant 'discovery' bottleneck is highly repairable via self-distillation offers a practical, high-impact post-training pathway for enhancing mathematical and logical reasoning in next-generation models.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
How to use it
- 1LLM mathematical diagnostics: Using the four-dimension framework to pinpoint model vulnerabilities in structural math understanding.
- 2Post-training & distillation: Employing primitive-guided self-distillation to repair and enhance open-source LLMs.
Limitations & caveats
- The effectiveness of the self-distillation framework still heavily depends on the quality of the mathematical primitives generated by the teacher model.
- Currently restricted to mathematics; whether the concept of primitives can generalize to other reasoning tasks like code generation remains to be tested.
Related

Ai2 Open-Sources AstaBrief: An 8B Scientific Report Generator 3.5x Faster than Claude
艾倫人工智慧研究所開源 AstaBrief:比 Claude 快 3.5 倍的 8B 科學報告生成模型
Allen Institute for AI (Ai2) has open-sourced AstaBrief 8B, a specialized model for scientific report generation that achieves a 3.5x speedup over proprietary pipelines while maintaining high citation accuracy.
GALA: Distilling 3D Gaussian Avatars into Linear Blendshapes for Real-Time Animation
GALA:用線性混合變形蒸餾技術實現 3D Gaussian 虛擬化身即時動畫
GALA distills complex neural decoding of 3D Gaussian avatars into lightweight linear blendshapes, reducing CPU animation costs by up to 1000x and enabling 60fps real-time performance on mobile devices.
ScholarCatalyst: A Benchmark for Testing AI's Intuition in Retrieving Inspiring Research Papers
ScholarCatalyst:評估 AI 是否擁有「科學家直覺」的學術文獻檢索基準
ScholarCatalyst is a novel benchmark featuring annotations from 184 lead authors to evaluate whether AI can retrieve key inspiring papers from past literature based only on an initial research question.