Most coverage of language model progress focuses on scale, data, or architecture. The quieter story, and arguably the more consequential one, is how the reward