Published collection
GRPO study shows very small gaps in native-language reasoning training
The findings can help multilingual reasoning models choose reinforcement learning training languages.
- Published entries
- 1
- Sources
- 1
- Date range
- 2026-08-19
- Primary labels
- Research · 论文 · Training Methods · GRPO 多语言研究 · Apple
Published evidence
Every entry keeps its summary and a path back to the source context.
GRPO study shows very small gaps in native-language reasoning training
Research from Apple Machine Learning Research examines GRPO performance in multilingual and non-English environments, covering multiple base models, training languages, and inference-language reward settings.
Why it mattersThe findings can help multilingual reasoning models choose reinforcement learning training languages.
Original source