NVIDIA said in a Hugging Face post on 7 October 2026 that systems built on its Nemotron 3 models reached gold-medal level at both the 2026 International Olympiad in Informatics (IOI) and the 2026 International Mathematical Olympiad (IMO).
At IOI 2026, a competitive-coding version of Nemotron 3 Ultra scored 535.4 out of 600, above the 361.12 gold threshold and the top human score of 498.27, in what NVIDIA calls an unofficial, unsupervised run that was not included in the official ranking. At IMO 2026, a proof system using three Nemotron 3 Ultra checkpoints scored 30 out of 42 against a gold threshold of 29, with its proofs graded by official IMO graders, the post says.
Business Pill · FINE-TUNING
A one-minute explainer of fine-tuning: taking a trained model and training it further on a narrower set of examples. It teaches the general idea only and says nothing about any company in this story.
The key insight: As we read it, NVIDIA’s own IOI 2025 figures put the gain in two places: post-training took Nano from 130 to 291 points, and the GenCorrect loop took it from 291 to 468, past the 438.3 gold threshold.
What NVIDIA Reported
The post sets out both results in one table. For IOI 2026 it names the system as Nemotron-3-Ultra-CC with supervised fine-tuning and GenCorrect, NVIDIA’s name for its iterative generate-evaluate-refine strategy. For IMO 2026 it names Nemotron 3 Ultra general, SFT and RL checkpoints run in a generate-verify-refine system.
The post states the conditions in two sentences: “The IOI result came from a live, prospective run under the same time, internet-access, and submission constraints as human contestants. It was an unofficial, unsupervised benchmark and was not included in the official IOI ranking.” It adds: “The IMO system’s submitted proofs were graded by official IMO graders.”
The IOI paper’s abstract goes one step further: “To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set.” That is NVIDIA’s researchers’ claim, made in their own paper.

How the Coding Model Was Built
For competitive programming, the post says NVIDIA curated 22,000 problems and generated synthetic reasoning traces to train two specialists. Nemotron-3-Nano-CC, with 30 billion total and 3 billion active parameters, received supervised fine-tuning and reinforcement learning. Nemotron-3-Ultra-CC, with 550 billion total and 55 billion active parameters, received supervised fine-tuning only.
On the IOI 2025 problem set, the post says Nano improved from 130 points before post-training to 280 after SFT and 291 after RL. With GenCorrect it reached 468 points, above the 438.3 gold threshold, and Ultra-CC reached 502 points with the same test-time strategy.
The post says SFT produced most of Nano’s gain, with RL adding a smaller but consistent improvement, and that one SFT epoch was enough for the larger Ultra model to outperform the fully post-trained Nano across IOI, ICPC and LiveCodeBench Pro. That finding, it says, guided the competition-specific Ultra-CC system used for IOI 2026.

How the Proof System Worked
For the IMO, the post says NVIDIA trained one specialist with SFT and another with RL, both starting from Nemotron 3 Ultra. The SFT corpus held 414,890 quality-filtered examples across 15,818 unique proof problems, covering proof generation, refinement, verification and meta-verification. The RL model was trained on 9,597 proof problems.
For each problem, the models generated candidate proofs, scored them, produced critiques and refined the most promising attempts, and a separate high-compute stage selected the final submission. The post says the system worked in natural language, “with no formal prover, external tools, or internet access.”
It scored 30 out of 42 points, the post says, including full credit on four of the six problems. It also says using the complementary SFT and RL checkpoints was more valuable than drawing more samples from one checkpoint.
What NVIDIA Released
The post lists what is public. A Nemotron Labs IMO 2026 collection on Hugging Face holds the SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench, a new benchmark of 200 olympiad-level problems. NVIDIA’s NeMo-Skills repository holds the IMO inference pipeline, prompts, submitted proofs and a quickstart.
For competitive programming, the post says the Nemotron-3-Ultra-CC model is available on Hugging Face, the IOI paper gives the training recipe and the GenCorrect method, and the IOI evaluation and inference pipeline is in NeMo-Skills.
NVIDIA frames the work as a recipe in four parts: start with a strong Nemotron base model, curate domain problems and reasoning traces, apply SFT and where useful RL, and pair the specialist with an inference loop that generates, evaluates and improves answers. The post says: “We did not need to build a new foundation model for every challenge.”
The Structural Read
The base model was reused, not rebuilt. The post says both specialists started from Nemotron 3 models, and that NVIDIA did not need to build a new foundation model for every challenge.
The inference loop carried much of the gain. By our arithmetic on NVIDIA’s IOI 2025 figures, GenCorrect added 177 points to Nano after post-training, while SFT and RL together added 161.
The two contests were scored on different terms. NVIDIA says the IOI run was unofficial and outside the official ranking, while the IMO proofs were graded by official IMO graders.
NVIDIA on Hugging Face, 7 October 2026
“The medals were not produced by fine-tuning alone, and they were not produced by brute-force sampling alone.”
Three Implications
TEAMS BUILDING ON OPEN MODELS The post says the IMO checkpoints, both training datasets, the inference pipeline, the prompts and what it calls a reproducible quickstart are published, along with the IOI model.
READERS OF AI BENCHMARK CLAIMS The post reports the two results on different footings: an unofficial IOI run outside the official ranking, and IMO proofs graded by official IMO graders.
COMPUTE BUDGETS The post calls the training and inference runs substantial and gives no figure, so the cost of reaching these scores is not stated.
The Business Engineer Lens
This story maps onto the Business Engineer framework The Open vs Closed Meta-Framework.
The framework starts from this line: “Every technology stack has layers. Some are scarce, some are abundant.”
As we read it, the post puts the opened layer in plain view: model weights, training data, inference code and a benchmark are published on Hugging Face and in NeMo-Skills, while the post describes the training and inference runs behind them only as substantial.
What Is Not Established
The post gives no figure for compute. It says only that “The training and inference runs were substantial”, with no GPU count, hours or cost.
The IOI score is NVIDIA’s own measurement of an unofficial run; the post says it was not part of the official IOI ranking. The IMO score rests on grading by official IMO graders, per the post. This publication has also covered OpenAI’s report that an internal model produced 722 math manuscripts.
The Bottom Line
NVIDIA’s 7 October post says fine-tuned Nemotron 3 systems scored 535.4 of 600 in an unofficial IOI 2026 run, above the top human score of 498.27, and 30 of 42 at IMO 2026, above the gold threshold of 29. It attributes both to the same pattern of a base model, curated data, post-training and a verify-and-refine loop, and has published the IMO checkpoints, data and code and the IOI model.
94,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
A note on sourcing. On 10 October 2026 we read NVIDIA’s 7 October post on Hugging Face in full, and the abstracts of the two papers it links on arXiv. The scores, conditions and training details are NVIDIA’s as stated there; we did not rerun any evaluation, and the IOI result is not part of the official IOI ranking. Nothing here is a forecast, and nothing here is financial or investment advice.
Sources: NVIDIA on Hugging Face: One Model Family, Two Gold-Level Results (7 Oct 2026) · IOI paper, arXiv 2609.02849 · IMO paper, arXiv 2609.10712








