NVIDIA has published a combined account of how it specialized Nemotron 3 models for two different olympiad-style challenges: competitive programming at the International Olympiad in Informatics and natural-language proof solving at the International Mathematical Olympiad. The October 7 release brings together models, datasets, papers and inference recipes for systems that NVIDIA reports crossed the gold-medal thresholds at both 2026 competitions.
For IOI, NVIDIA used supervised fine-tuning and reinforcement learning on specialist Nemotron models trained from 22,000 curated programming problems and synthetic reasoning traces. The larger Nemotron-3-Ultra-CC has 550 billion total parameters with 55 billion active parameters and received supervised fine-tuning. NVIDIA paired the specialists with GenCorrect, a test-time loop that generates solutions, evaluates them and iteratively refines candidates. In a prospective IOI 2026 run, the company reports a score of 535.4 out of 600, above the official gold threshold of 361.12 and the top human score of 498.27.
That IOI result requires an important qualification. NVIDIA says the run used the same time, internet-access and submission constraints as human contestants, but it was an unofficial and unsupervised evaluation and was not part of the official IOI ranking. The associated research paper provides the methodology and reports the same score, while independent coverage has correctly kept the result framed as an NVIDIA research claim rather than an officially ranked AI contestant.
The IMO system used a different specialization recipe. NVIDIA trained supervised and reinforcement-learning checkpoints from Nemotron 3 Ultra and combined them with the generally available model in a generate-verify-refine pipeline. Models produced candidate proofs, scored and critiqued them, refined promising answers and then used a higher-compute selection stage. NVIDIA says the submitted natural-language proofs were graded by official IMO graders and received 30 out of 42 points, one point above the 2026 gold threshold, with full credit on four of six problems.
The release is more significant than a pair of headline scores because NVIDIA is exposing much of the machinery behind them. The Nemotron Labs IMO collection includes specialist checkpoints, training datasets and a 200-problem benchmark, while NeMo-Skills contains the inference pipeline, prompts and submitted proofs. The competitive-programming model and its methodology are also available through Hugging Face and the associated paper. That makes the work more inspectable than a closed demonstration, although inspectability is not the same as independent replication.
The broader research claim is that a strong foundation model can be turned into a domain specialist using conventional post-training plus structured test-time computation, rather than requiring a new foundation model for every difficult domain. The evidence is substantial enough to merit attention, but the two competitions have different verification strength. The IMO grading has an official external component; the IOI score remains a prospective, vendor-run benchmark that was outside the official ranking. Independent reproduction of the full pipelines would provide the strongest next test.