---
id: 20260913-T0-03
title: "Nemotron奥数金牌配方公开：后训练和推理设计如何影响证明生成"
title_en: "Open Recipe for IMO Gold: How Post-Training and Inference Design Affect Nemotron's Math Proofs"
url: https://ai.daily.yangsir.net/daily/20260913-T0-03
issue_date: 2026-09-13
publish_date: 2026-09-12T04:00:00.000Z
category: research
source_name: "arXiv cs.AI"
source_url: https://arxiv.org/abs/2609.10712
---

# Nemotron奥数金牌配方公开：后训练和推理设计如何影响证明生成

arXiv论文2609.10712研究模型后训练和测试时推理设计如何影响高难度奥数自然语言证明生成。团队从Nemotron 3 Ultra出发，用监督微调和强化学习训练了两个专家checkpoint，并对比不同推理策略的效果。研究给出了可复现的训练配方，目标是让模型在奥林匹克数学级别的证明任务上达到金牌水平。对关注LLM数学推理能力的研究者和开发者来说，这份配方提供了具体可操作的训练路径。

## English Version

**Open Recipe for IMO Gold: How Post-Training and Inference Design Affect Nemotron's Math Proofs**

arXiv paper 2609.10712 studies how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, the team trained two specialist checkpoints using supervised fine-tuning and reinforcement learning, comparing different inference strategies. The work provides a reproducible training recipe aiming for gold-medal-level performance on Olympiad math proof tasks. For researchers and developers tracking LLM math reasoning, it offers a concrete, actionable training path.

---

**来源**：[arXiv cs.AI](https://arxiv.org/abs/2609.10712)

**详情页**：https://ai.daily.yangsir.net/daily/20260913-T0-03

---

*智语观潮 · Daily — https://ai.daily.yangsir.net/llms.txt*