Skip to content

Published collection

Grok and Gemini score higher at medium effort in deepswe

When choosing a reasoning tier, developers should not look only at the level; they also need to verify score, cost, and speed.

Updated
Editorial
Frontline Lab
Published entries
1
Sources
1
Date range
2026-08-18
Primary labels
Models · 模型 · Model evaluations · Grok 和 Gemini · Google · xAI

Published evidence

Every entry keeps its summary and a path back to the source context.

X

Grok and Gemini score higher at medium effort in deepswe

The post says the latest grok and Gemini show an inverted reasoning-tier result in the deepswe benchmark: medium effort scores higher than high/xhigh, while also being cheaper and faster.

Why it mattersWhen choosing a reasoning tier, developers should not look only at the level; they also need to verify score, cost, and speed.

Original source