One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

2026-10-10 · Hugging Face

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

Overview

The International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) test different skills. IOI requires algorithms and code that pass hidden tests under strict time and submission limits. IMO demands rigorous natural-language proofs. NVIDIA's teams used Nemotron 3 as a foundation, applying supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to achieve gold-medal level at both IMO 2026 and IOI 2026.

Competition Results

| Competition | Nemotron Specialization | Result |

|-------------|------------------------|--------|

| IOI 2026 | Nemotron-3-Ultra-CC with SFT and GenCorrect | 535.4/600, above the 361.12 gold threshold and top human score of 498.27 |

| IMO 2026 | Nemotron 3 Ultra general, SFT, and RL checkpoints in generate-verify-refine system | 30/42, above the official gold threshold of 29 |

The IOI result came from a live, prospective run under the same time, internet-access, and submission constraints as human contestants. It was an unofficial, unsupervised benchmark not included in official IOI ranking. The IMO system's submitted proofs were graded by official IMO graders.

A Reusable Specialization Recipe

"Easy to fine-tune" should mean more than making a checkpoint trainable. It should mean that a capable foundation model can be adapted to a demanding domain with a clear, reusable recipe. Across both projects, the recipe had four parts:

1. Start with a strong Nemotron base model

2. Curate domain-specific problems and high-quality reasoning traces

3. Apply standard post-training methods such as SFT and, where useful, RL

4. Pair the specialist model with an inference loop that generates, evaluates, and improves candidate answers

The training and inference runs were substantial, but the underlying approach is familiar and reproducible. The team did not need to build a new foundation model for every challenge—they specialized Nemotron for the task.

From General Coding Ability to IOI Gold

For competitive programming, the team curated 22,000 problems and generated synthetic reasoning traces to train two specialists:

  • Nemotron-3-Nano-CC: 30 billion total parameters, 3 billion active parameters, received both SFT and RL
  • Nemotron-3-Ultra-CC: 550 billion total parameters, 55 billion active parameters, received SFT

The progression on IOI 2025 demonstrates the value of specialization:

  • Nano improved from 130 points before post-training to 280 after SFT and 291 after RL
  • With GenCorrect, the iterative generate-evaluate-refine strategy, it reached 468 points, crossing the gold threshold of 438.3
  • Ultra-CC reached 502 points with the same test-time strategy

Experiments showed that adaptation does not have to look the same at every scale. SFT produced most of Nano's gain, with RL adding a smaller but consistent improvement. For the stronger Ultra model, one SFT epoch was enough to outperform the fully post-trained Nano model across IOI, ICPC, and LiveCodeBench Pro. This finding guided the competition-specific Ultra-CC system used for IOI 2026, which scored 535.4 out of 600.

Teaching Nemotron to Prove, Check, and Revise

The IMO project applied the same idea to olympiad mathematics. Starting from Nemotron 3 Ultra, the team trained one specialist with SFT and another with RL.

  • SFT corpus: 414,890 quality-filtered examples across 15,818 unique proof problems. The data covered proof generation, refinement, verification, and meta-verification, teaching the model to construct arguments, identify gaps, respond to critiques, and judge whether a proof was complete.
  • RL model: Trained on 9,597 proof problems selected near the model's capability frontier.

Both post-trained checkpoints outperformed the general-availability model in development experiments. The SFT checkpoint was strongest in the first search round, while the RL checkpoint achieved the best overall single-checkpoint result. Their strengths were complementary, so the final system used both specialists alongside the general model.

For each IMO problem, the models generated candidate proofs, scored them, produced critiques, and refined the most promising attempts. A separate high-compute stage selected the final submission. The entire system worked in natural language, with no formal prover, external tools, or internet access. It scored 30 out of 42 points, including full credit on four of six problems, exceeding the official gold-medal threshold.

Fine-Tuning and Test-Time Compute Work Together

The team's earlier IOI 2025 Hugging Face post showed how test-time compute can push open-weight models to gold-level performance. The new results add an important piece: better specialization gives the inference system better candidates, better critics, and better refinements.

  • At IOI, GenCorrect turned fine-tuning gains into larger improvements over multiple feedback rounds
  • At IMO, using complementary SFT and RL checkpoints was more valuable than simply drawing more samples from one checkpoint

The medals were not produced by fine-tuning alone, nor by brute-force sampling alone. They came from co-designing the model, the data, and the inference loop.

Open Models, Data, and Recipes on Hugging Face

The team wants these results to be useful beyond the competitions. The Nemotron Labs IMO 2026 collection brings together open models, data, and recipes on Hugging Face for the community to build upon.

Source