Rate My Model
← Back to leaderboard

Chart gallery

See the same DeepSWE leaderboard through several chart types. Use the view that makes your next model decision easiest.

Snapshot: 2026-08-13T16:11:55.708636+00:00

Showing 45 leaderboard rows. Each row is a model plus its reasoning effort.

Efficiency frontier

Look for points toward the upper-left: higher pass rate with less mean cost is the strongest trade-off.

Providers
  • Anthropic
  • OpenAI
  • Moonshot
  • xAI
  • Luna
  • DeepSeek
  • Alibaba
  • Meta
  • Zhipu
  • Google

Chart loads in the browser.

Pass rate by model

Compare quality directly, with the highest-scoring model-and-effort rows at the top.

Providers
  • Anthropic
  • OpenAI
  • Moonshot
  • xAI
  • Luna
  • DeepSeek
  • Alibaba
  • Meta
  • Zhipu
  • Google

Chart loads in the browser.

Cost by model

Cheapest rows rise to the top, making low-cost alternatives easy to spot before comparing their quality.

Providers
  • Anthropic
  • OpenAI
  • Moonshot
  • xAI
  • Luna
  • DeepSeek
  • Alibaba
  • Meta
  • Zhipu
  • Google

Chart loads in the browser.

Output tokens vs pass rate

Longer outputs can signal thoroughness or verbosity; compare that trade-off against quality rather than assuming more tokens are better.

Providers
  • Anthropic
  • OpenAI
  • Moonshot
  • xAI
  • Luna
  • DeepSeek
  • Alibaba
  • Meta
  • Zhipu
  • Google

Chart loads in the browser.

Cost per pass-rate point

This value ranking divides mean cost by pass-rate percentage points; shorter bars are the more economical quality wins.

Providers
  • Anthropic
  • OpenAI
  • Moonshot
  • xAI
  • Luna
  • DeepSeek
  • Alibaba
  • Meta
  • Zhipu
  • Google

Chart loads in the browser.