Nvidia has announced that specialized versions of its Nemotron 3 model family achieved gold-medal-level performance in both the International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) for 2026. The results, detailed in a Hugging Face blog post, demonstrate that a single foundation model can be adapted via fine-tuning and inference strategies to excel in two distinct, highly demanding competitive domains.
What Happened
For the IOI, Nvidia reported that Nemotron-3-Ultra-CC, trained with supervised fine-tuning (SFT) and a generate-evaluate-refine strategy called GenCorrect, scored 535.4 out of 600. This score exceeded the official gold-medal threshold of 361.12 and surpassed the top human contestant score of 498.27. The evaluation was conducted as a live, prospective run under the same time and submission constraints as human participants, though Nvidia noted it was an unofficial, unsupervised benchmark not included in the official IOI ranking.
In the IMO, a system combining Nemotron 3 Ultra general, SFT, and reinforcement learning (RL) checkpoints achieved a score of 30 out of 42. This result exceeded the official gold-medal threshold of 29 points. The proofs submitted by the system were graded by official IMO graders. The model worked entirely in natural language without external tools, formal provers, or internet access, achieving full credit on four of the six problems.
The company described a reusable specialization recipe consisting of four steps: starting with a strong Nemotron base model, curating domain-specific problems and reasoning traces, applying standard post-training methods like SFT and RL, and pairing the specialist model with an inference loop that generates, evaluates, and improves candidate answers.
Why It Matters
These results highlight a shift in how frontier models are optimized for specific high-stakes tasks. Nvidia emphasized that the gold-level outcomes were not produced by fine-tuning alone or by brute-force sampling alone, but by co-designing the model, the data, and the inference loop. For developers, this suggests that test-time compute strategies, such as GenCorrect, can significantly amplify the gains from specialized fine-tuning.
The data also reveals nuances in model scaling. For the smaller Nemotron-3-Nano-CC model (30 billion total parameters), RL provided a smaller but consistent improvement over SFT. However, for the larger Nemotron-3-Ultra model (550 billion total parameters), a single epoch of SFT was sufficient to outperform the fully post-trained Nano model across multiple benchmarks, including IOI, ICPC, and LiveCodeBench Pro. This finding guided the competition-specific configuration for the IOI 2026 entry.
Nvidia has released the Nemotron Labs IMO 2026 collection on Hugging Face, which includes the SFT and RL checkpoints, training datasets, and a new benchmark called Nemotron-IMO-Bench with 200 olympiad-level problems. The Nemotron-3-Ultra-CC model and the IOI training recipe are also available, alongside reproducible inference pipelines in the NeMo-Skills repository.
The Bottom Line
Nvidia’s Nemotron 3 family has demonstrated the ability to reach gold-medal performance in both IOI and IMO 2026 through targeted fine-tuning and advanced inference workflows. By releasing the models, data, and recipes, the company aims to provide the community with a reproducible framework for adapting foundation models to specialized, high-difficulty domains.