Training and post-training
Reasoning and test-time compute
Post-training that teaches a model to produce a long reasoning trace before answering, and to spend more compute at inference when a problem is harder. Reinforcement learning with verifiable rewards — where a checkable answer supplies the training signal — is what made the approach reproducible outside closed labs. The trade is worth stating plainly: accuracy bought with generation length rather than with a larger model, and paid for in latency and serving cost.
Key terms
- reasoning traces
- verifiable rewards
- inference budget
- chain of thought
Connected across the project
Builds on
Related technology
Related papers
Historical context
Primary sources
This is a curated technical map, not a claim of comprehensive coverage.