Double Gold: Specializing NVIDIA Nemotron for IOI and IMO 2026
雙金入袋!NVIDIA 透過微調 Nemotron 征服 2026 國際奧賽數理與資訊殿堂

NVIDIA demonstrated how a four-step specialization recipe can transform Nemotron-3 base models into domain experts. At IOI 2026, the 550B-parameter Nemotron-3-Ultra-CC scored 535.4/600, surpassing both the gold threshold and the top human score. At IMO 2026, the specialized system scored 30/42, beating the gold threshold using pure natural language reasoning without external provers. By co-designing model tuning, high-quality data, and test-time search loops (GenCorrect), they achieved frontier-level problem solving, now fully open-sourced.
Key points
The Specialization Recipe
Rather than building new foundations, the team combined strong base models with curated data, post-training (SFT/RL), and a 'generate-evaluate-refine' inference loop.
Tuning Dynamics at Scale
For the 30B Nano model, combining SFT and RL yielded steady gains. For the 550B Ultra model, just one epoch of SFT outperformed a fully-tuned Nano model.
Natural Language Reasoning
The IMO system used no formal provers, external tools, or web access, relying entirely on natural language proof generation, critiquing, and refining.
Test-Time Co-Design
Gold results relied on combining capable specialized models with structured test-time search and verification loops (such as GenCorrect) during inference.
How it works
| IOI 2026 (程式/資訊) | IMO 2026 (數學) | |
|---|---|---|
| Specialized Models | Nemotron-3-Ultra-CC (550B SFT) | Nemotron 3 Ultra (SFT 與 RL 組合) |
| Score achieved | 535.4 / 600 (金牌門檻 361.12) | 30 / 42 (金牌門檻 29) |
| Key Strategy | GenCorrect 多輪迭代修正迴圈 | 生成-驗證-重構 多模型協同審查 |
| External Tools | 有(程式編譯與回饋) | 無(完全基於自然語言推理) |
Why it matters
This achievement proves that open-weight base models can reach human-olympiad-level performance without retraining from scratch. By using a structured, reproducible specialization recipe and test-time search, enterprises and researchers can replicate this success in other complex, high-reasoning domains like software engineering and scientific research.
Who it affects
- AI Developer
- AI Researcher
- Product Manager
- Enterprise Leader
How to use it
- 1Advanced Coding & Debugging: Employing multi-round 'generate-evaluate-refine' (GenCorrect) workflows for automated software engineering and complex algorithm synthesis.
- 2Complex Logic & Academic Proofing: Generating and self-verifying rigorous arguments in high-precision fields like math, legal analysis, or finance using SFT and meta-verification.
Limitations & caveats
- Unofficial & Unsupervised Runs: The IOI benchmark was a simulated prospective run under competition constraints, not an official supervised entry.
- High Test-Time Compute Overhead: Achieving gold-level performance relies heavily on multi-round search and refinement, increasing latency and operational compute costs during inference.
Related
Building Persistent 3D Object Memory: How Ledger Tracks Objects from Egocentric Videos
打造過目不忘的 3D 空間記憶:Ledger 如何透過第一人稱影片追蹤隱形物體
Researchers introduce Ledger, a framework that builds a persistent 3D object memory from egocentric videos, significantly improving spatial question-answering accuracy for embodied agents.
Decoupling Exploration from Optimization: How ExpDis Boosts LLM Reasoning and Solution Diversity
探索與優化解耦:全新強化學習框架 ExpDis 提升大語言模型的推理多元性
The ExpDis framework decouples exploration from optimization in RLVR. By training explorers with novelty bonuses and distilling filtered trajectories into a student model, it prevents model degradation while fostering diverse reasoning.
Clipped Decentralized SGD: Achieving Optimal Convergence and Linear Speed-Up Under Heavy-Tailed Noise
去中心化 SGD 克服重尾雜訊:梯度裁剪如何實現最佳收斂與線性加速
This study proves that clipped decentralized SGD (DSGD) achieves order-optimal convergence rates and linear speed-up under heavy-tailed noise for non-convex optimization.