Frontline Lab summary and source
Microsoft Research says multilingual Agent evaluations often compare only final answers and discard trajectories; the paper uses action policy as the metric, covering 8 models, 6 parallel benchmarks, 41 languages, and 2.38M rollouts, and analyzes five types of confounding factors.
This brief preserves the original source so the summary and editorial context can be checked independently.
Source attributionX · @dair_ai
Open the original source