Frontline Lab summary and source
The post says the latest grok and Gemini show an inverted reasoning-tier result in the deepswe benchmark: medium effort scores higher than high/xhigh, while also being cheaper and faster.
This brief preserves the original source so the summary and editorial context can be checked independently.
Source attributionX · @_kaichen
Open the original source