English
← Back to AI Technology

Training and post-training

Reasoning and test-time compute

Post-training that teaches a model to produce a long reasoning trace before answering, and to spend more compute at inference when a problem is harder. Reinforcement learning with verifiable rewards — where a checkable answer supplies the training signal — is what made the approach reproducible outside closed labs. The trade is worth stating plainly: accuracy bought with generation length rather than with a larger model, and paid for in latency and serving cost.

Reference period
2024 / 2025
Last reviewed

Key terms

  • reasoning traces
  • verifiable rewards
  • inference budget
  • chain of thought

Connected across the project

Primary sources

  1. OpenAI o1 System Card
  2. DeepSeek-R1
  3. s1: Simple Test-Time Scaling

This is a curated technical map, not a claim of comprehensive coverage.