DeepSeek — Field Evidence
paperUnverified
DeepSeek · DeepSeek · adoption
Field note
On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards — Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1
collected 2026-07-28original 2026-07-25
Does this shift the US–China race?
Be the first to call it
Impact on the index
Adoption · minor
China +20
Directional contribution — before recency decay and per-type diminishing returns. How it works →
Related Front
US vs ChinaLikely
U.S. FrontiervsChina Open-Weight
Likely — capability gap narrowing on common tasks
Letters from the Front
Be the first to file a report.