Aivora
Hugging FaceAI ResearchIntermediate

Double Gold: Specializing NVIDIA Nemotron for IOI and IMO 2026

雙金入袋!NVIDIA 透過微調 Nemotron 征服 2026 國際奧賽數理與資訊殿堂

2 min read
Double Gold: Specializing NVIDIA Nemotron for IOI and IMO 2026
The 30-second version

NVIDIA demonstrated how a four-step specialization recipe can transform Nemotron-3 base models into domain experts. At IOI 2026, the 550B-parameter Nemotron-3-Ultra-CC scored 535.4/600, surpassing both the gold threshold and the top human score. At IMO 2026, the specialized system scored 30/42, beating the gold threshold using pure natural language reasoning without external provers. By co-designing model tuning, high-quality data, and test-time search loops (GenCorrect), they achieved frontier-level problem solving, now fully open-sourced.

Key points

01

The Specialization Recipe

Rather than building new foundations, the team combined strong base models with curated data, post-training (SFT/RL), and a 'generate-evaluate-refine' inference loop.

02

Tuning Dynamics at Scale

For the 30B Nano model, combining SFT and RL yielded steady gains. For the 550B Ultra model, just one epoch of SFT outperformed a fully-tuned Nano model.

03

Natural Language Reasoning

The IMO system used no formal provers, external tools, or web access, relying entirely on natural language proof generation, critiquing, and refining.

04

Test-Time Co-Design

Gold results relied on combining capable specialized models with structured test-time search and verification loops (such as GenCorrect) during inference.

How it works

Comparison of IOI vs IMO 2026 Specialized Systems
IOI 2026 (程式/資訊)IMO 2026 (數學)
Specialized ModelsNemotron-3-Ultra-CC (550B SFT)Nemotron 3 Ultra (SFT 與 RL 組合)
Score achieved535.4 / 600 (金牌門檻 361.12)30 / 42 (金牌門檻 29)
Key StrategyGenCorrect 多輪迭代修正迴圈生成-驗證-重構 多模型協同審查
External Tools有(程式編譯與回饋)無(完全基於自然語言推理)

Why it matters

This achievement proves that open-weight base models can reach human-olympiad-level performance without retraining from scratch. By using a structured, reproducible specialization recipe and test-time search, enterprises and researchers can replicate this success in other complex, high-reasoning domains like software engineering and scientific research.

Who it affects

  • AI Developer
  • AI Researcher
  • Product Manager
  • Enterprise Leader

How to use it

  1. 1Advanced Coding & Debugging: Employing multi-round 'generate-evaluate-refine' (GenCorrect) workflows for automated software engineering and complex algorithm synthesis.
  2. 2Complex Logic & Academic Proofing: Generating and self-verifying rigorous arguments in high-precision fields like math, legal analysis, or finance using SFT and meta-verification.

Limitations & caveats

  • Unofficial & Unsupervised Runs: The IOI benchmark was a simulated prospective run under competition constraints, not an official supervised entry.
  • High Test-Time Compute Overhead: Achieving gold-level performance relies heavily on multi-round search and refinement, increasing latency and operational compute costs during inference.

Related