Stories about PPO
1 related stories
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
AI InsightRLVR improves pass@1 but contracts the policy's solution space, with coverage dropping by up to 67% on Countdown. Unlike prior focus on accuracy gains, this work quantifies diversity loss at trajectory entrances, indicating diminishing test-time scaling returns stem from restricted access rather than execution failures.Key TakeawayFirst quantification of RLVR causing up to 67% solution-space coverage drop.Why It MattersReveals RLVR's hidden cost, challenging the simple equation of verifiable rewards with better models, and impacting training and test-time scaling design.Who's Affected- AI ResearchersMust rebalance RLVR accuracy gains against solution-space diversity loss and explore diversity-preserving reward designs.
- DevelopersWhen fine-tuning with RLVR, evaluate coverage on long-tail problems to avoid failure on rare but valid reasoning paths.
- AI ResearchersNew explanation for limited test-time scaling gains may drive research on restoring solution-space access.
What's NextWatch for training methods that mitigate solution-space contraction and whether findings generalize to more complex reasoning tasks.Importance 75/100