Anthropic
Grade My Model / Rate My Model
Explore DeepSWE model performance, pricing, and comparisons in one place.
Leaderboard snapshot: 2026-08-13T16:11:55.708636+00:00
Model selector
Choose the model and reasoning effort rows to compare.
Anthropic
Claude Opus
Anthropic
Claude Sonnet
DeepSeek
Deepseek V Flash
DeepSeek
Deepseek V Pro
Gemini 3 Flash
Zhipu
Glm
Luna
Gpt 5 Luna
OpenAI
Gpt 5 Sol
OpenAI
Gpt 5 Terra
xAI
Grok 4
Moonshot
Kimi K
Moonshot
Kimi K Code
Meta
Muse Spark
Alibaba
Qwen Max
Efficiency frontier
Compare DeepSWE pass rate with mean cost. The top-left corner is the sweet spot: higher performance for less money.
Chart loads in the browser.
Leaderboard table
15 selected models · sort by performance, price, or speed
| Model | |||
|---|---|---|---|
claude-opus-5 Anthropic max | 73.6% ±3.9 FilledUncertainEmpty | $11.84 $0.16 / pass point lower is cheaper | 113k 90.5 steps |
gpt-5-6-sol OpenAI max | 72.7% ±2.8 | $8.39 $0.12 / pass point lower is cheaper | 59k 53 steps |
claude-fable-5 Anthropic xhigh | 69.9% ±3.2 | $13.41 $0.19 / pass point lower is cheaper | 76k 61.5 steps |
gpt-5-6-terra OpenAI max | 69.6% ±2.6 | $4.95 $0.07 / pass point lower is cheaper | 71k 71 steps |
kimi-k3 Moonshot max | 68.5% ±4.5 | $4.65 $0.07 / pass point lower is cheaper | 75k 88 steps |
grok-4-6 xAI medium | 67.5% ±2.3 | $3.45 $0.05 / pass point lower is cheaper | 49k 66 steps |
gpt-5-6-luna Luna max | 67.2% ±4.0 | $3.03 $0.05 / pass point lower is cheaper | 70k 92.5 steps |
gemini-3-7-flash Google medium | 65.5% ±3.1 | $2.03 $0.03 / pass point lower is cheaper | 88k 108 steps |
deepseek-v4-pro DeepSeek max | 62.8% ±6.3 | $0.24 $0.0038 / pass point lower is cheaper | 101k 146 steps |
qwen3-8-max Alibaba xhigh | 57.5% ±2.7 | $3.73 $0.06 / pass point lower is cheaper | 90k 102 steps |
muse-spark-1-2 Meta xhigh | 54.9% ±2.1 | $3.70 $0.07 / pass point lower is cheaper | 81k 94 steps |
claude-sonnet-5 Anthropic max | 53.8% ±4.2 | $26.40 $0.49 / pass point lower is cheaper | 204k 260 steps |
deepseek-v4-flash DeepSeek max | 53.3% ±3.6 | $0.10 $0.0019 / pass point lower is cheaper | 104k 148 steps |
glm-5-2 Zhipu max | 43.8% ±1.7 | $3.92 $0.09 / pass point lower is cheaper | 76k 123 steps |
kimi-k2-7-code Moonshot default | 30.5% ±0.5 | $2.82 $0.09 / pass point lower is cheaper | 54k 139 steps |
Auto value comparison
15 selected models · alternatives come from the full visible leaderboard
Selected model
claude-opus-5
max · 73.6% · $11.84
→ Better value:
claude-opus-5
high · 72.8% · $6.08
Save $5.76 (48.7%) at ~same score
Selected model
gpt-5-6-sol
max · 72.7% · $8.39
→ Better value:
gpt-5-6-sol
xhigh · 70.7% · $4.70
Save $3.68 (43.9%) at ~same score
Selected model
claude-fable-5
xhigh · 69.9% · $13.41
→ Better value:
claude-opus-5
medium · 68.9% · $3.29
Save $10.12 (75.5%) at ~same score
Selected model
gpt-5-6-terra
max · 69.6% · $4.95
→ Better value:
claude-opus-5
medium · 68.9% · $3.29
Save $1.66 (33.5%) at ~same score
Selected model
kimi-k3
max · 68.5% · $4.65
→ Better value:
gpt-5-6-luna
max · 67.2% · $3.03
Save $1.63 (34.9%) at ~same score
Selected model
grok-4-6
medium · 67.5% · $3.45
→ Better value:
gemini-3-7-flash
medium · 65.5% · $2.03
Save $1.42 (41.3%) at ~same score
Selected model
gpt-5-6-luna
max · 67.2% · $3.03
→ Better value:
gemini-3-7-flash
medium · 65.5% · $2.03
Save $1.00 (33.1%) at ~same score
Selected model
gemini-3-7-flash
medium · 65.5% · $2.03
→ Better value:
No cheaper equal-score option on the board
Selected model
deepseek-v4-pro
max · 62.8% · $0.24
→ Better value:
No cheaper equal-score option on the board
Selected model
Higher score · lower costqwen3-8-max
xhigh · 57.5% · $3.73
→ Better value:
deepseek-v4-pro
max · 62.8% · $0.24
Save $3.49 (93.5%) at ~same score
Selected model
muse-spark-1-2
xhigh · 54.9% · $3.70
→ Better value:
deepseek-v4-flash
max · 53.3% · $0.10
Save $3.60 (97.3%) at ~same score
Selected model
claude-sonnet-5
max · 53.8% · $26.40
→ Better value:
deepseek-v4-flash
max · 53.3% · $0.10
Save $26.30 (99.6%) at ~same score
Selected model
deepseek-v4-flash
max · 53.3% · $0.10
→ Better value:
No cheaper equal-score option on the board
Selected model
Higher score · lower costglm-5-2
max · 43.8% · $3.92
→ Better value:
deepseek-v4-flash
max · 53.3% · $0.10
Save $3.82 (97.4%) at ~same score
Selected model
Higher score · lower costkimi-k2-7-code
default · 30.5% · $2.82
→ Better value:
deepseek-v4-flash
max · 53.3% · $0.10
Save $2.72 (96.4%) at ~same score
Luna value verdict
Luna models are compared with peers within three percentage points of their score.
gpt-5-6-luna
Not cheapest in band64.2–70.2% band · 67.2% · $3.03
gemini-3-7-flash is $1.00 cheaper at 65.5% (1.7 points lower than Luna).
Cheapest in band: gemini-3-7-flash at $2.03
gpt-5-6-luna
Best in class53.9–59.9% band · 56.9% · $1.54
Luna is the cheapest model in this performance band.
Cheapest in band: gpt-5-6-luna at $1.54
gpt-5-6-luna
Best in class41.2–47.2% band · 44.2% · $0.78
Luna is the cheapest model in this performance band.
Cheapest in band: gpt-5-6-luna at $0.78
gpt-5-6-luna
Best in class8.3–14.3% band · 11.3% · $0.22
Luna is the only model in this performance band.
Cheapest in band: gpt-5-6-luna at $0.22
gpt-5-6-luna
Best in class0.0–4.5% band · 1.5% · $0.07
Luna is the only model in this performance band.
Cheapest in band: gpt-5-6-luna at $0.07