DeepSWE 长周期编程 Agent 基准榜单

113
评测任务
91
仓库
5
编程语言
28
模型
数据同步于 2026-09-04 08:20:38
每个任务平均花费的美元金额。每个任务模型平均输出的 token 数量,token 越多通常代表推理链/回复越长。每个任务 agent 平均执行的操作步数(读文件、跑测试、改代码等各算一步),步数越多代表折腾的轮次越多。
0%20%40%60%80%$0$5$10$15每任务平均成本DeepSWE scoremost efficient ↗gpt-6-astraxhighgemini-3.8-flashhighclaude-opus-5maxgpt-5.6-solmaxclaude-fable-5xhighgpt-5.6-terramaxglm-5.3maxkimi-k3maxgrok-4.6mediumgpt-5.6-lunamaxgpt-5.5xhighgemini-3.7-flashmediumglm-5.3-flashmaxdeepseek-v4-promaxclaude-opus-4.8maxqwen3.8-maxxhighmuse-spark-1.2xhighclaude-sonnet-5xhigh
把每个模型在每一种思考程度下的成绩都列出来,同一模型会出现多行。每个模型会在多种思考程度下各跑一次,这里只保留每个模型 pass@1 分数最高的那一条,一个模型一行。
#
模型
PASS@1给模型 1 次机会做任务,不能重试。数字是能一次做对的任务比例,± 后面是多次评测统计出的置信区间,不是误差范围。
平均成本
输出 TOKEN
步数
1
gpt-6-astra[xhigh]
74% ±3%
$6.52
30k
29
2
gemini-3.8-flash[high]
74% ±1%
$2.36
143k
166
3
claude-opus-5[max]
74% ±4%
$11.84
118k
99
4
gpt-6-astra[high]
73% ±3%
$5.72
27k
27
5
gpt-6-astra[max]
73% ±1%
$12.37
61k
28
6
claude-opus-5[xhigh]
73% ±3%
$9.07
92k
89
7
claude-opus-5[high]
73% ±2%
$6.08
64k
73
8
gpt-6-astra[medium]
73% ±3%
$4.38
20k
26
9
gpt-5.6-sol[max]
73% ±3%
$6.46
60k
61
10
gemini-3.8-flash[medium]
71% ±2%
$1.97
125k
147
11
gpt-5.6-sol[xhigh]
71% ±1%
$3.60
41k
44
12
claude-fable-5[xhigh]
70% ±3%
$13.41
80k
68
13
claude-fable-5[max]
70% ±4%
$21.63
119k
88
14
gpt-5.6-terra[max]
70% ±3%
$3.96
72k
76
15
gpt-5.6-sol[high]
69% ±1%
$2.66
28k
37
16
glm-5.3[max]
69% ±3%
$3.99
80k
124
17
claude-opus-5[medium]
69% ±1%
$3.29
37k
52
18
claude-fable-5[high]
69% ±1%
$9.18
57k
59
19
kimi-k3[max]
69% ±5%
$4.65
81k
98
20
grok-4.6[medium]
67% ±2%
$3.45
50k
70
21
gpt-5.6-luna[max]
67% ±4%
$0.61
73k
102
22
gpt-5.5[xhigh]
67% ±6%
$7.23
46k
82
23
gpt-6-astra[low]
67% ±1%
$2.19
11k
20
24
grok-4.6[xhigh]
67% ±2%
$5.50
71k
87
25
gemini-3.7-flash[medium]
65% ±3%
$2.03
94k
117
26
claude-fable-5[medium]
65% ±4%
$6.09
40k
48
27
gemini-3.7-flash[high]
65% ±2%
$2.18
107k
125
28
grok-4.6[high]
65% ±2%
$4.38
61k
79
29
gpt-5.5[high]
64% ±3%
$5.10
31k
62
30
glm-5.3-flash[max]
63% ±4%
$0.24
73k
123
31
deepseek-v4-pro[max]
63% ±6%
$1.67
106k
155
32
gpt-5.6-sol[medium]
61% ±2%
$1.42
18k
31
33
gpt-5.6-terra[xhigh]
60% ±2%
$1.70
40k
43
34
claude-fable-5[low]
60% ±3%
$3.76
25k
38
35
claude-opus-4.8[max]
59% ±2%
$13.22
135k
120
36
claude-opus-5[low]
58% ±2%
$1.66
20k
36
37
qwen3.8-max[xhigh]
57% ±3%
$3.73
95k
111
38
gpt-5.6-luna[xhigh]
57% ±2%
$0.31
45k
71
39
muse-spark-1.2[xhigh]
55% ±2%
$3.70
99k
101
40
claude-opus-4.8[xhigh]
54% ±4%
$8.01
86k
95
41
gpt-5.5[medium]
54% ±3%
$2.75
20k
46
42
claude-sonnet-5[max]
54% ±4%
$26.40
214k
268
43
gemini-3.7-flash[low]
54% ±3%
$1.83
73k
130
44
gpt-5.6-terra[high]
54% ±4%
$0.91
22k
34
45
grok-4.5[high]
54% ±2%
$2.42
36k
61
46
deepseek-v4-flash[max]
53% ±4%
$0.46
108k
153
47
muse-spark-1.1[xhigh]
53% ±3%
$2.36
74k
96
48
claude-opus-4.8[high]
52% ±5%
$4.28
50k
73
49
gpt-5.4[xhigh]
52% ±2%
$5.65
71k
70
50
claude-sonnet-5[xhigh]
50% ±3%
$11.89
121k
186
51
claude-opus-4.8[medium]
49% ±2%
$3.44
41k
66
52
claude-sonnet-5[high]
48% ±5%
$7.43
87k
147
53
gemini-3.6-flash[high]
47% ±4%
$2.21
96k
117
54
gpt-5.6-sol[low]
45% ±2%
$0.82
11k
23
55
gpt-5.6-luna[high]
44% ±3%
$0.16
26k
49
56
glm-5.2[max]
44% ±2%
$3.92
78k
129
57
grok-4.6[low]
42% ±2%
$1.04
16k
44
58
claude-opus-4.8[low]
41% ±1%
$2.29
29k
54
59
claude-sonnet-5[medium]
40% ±3%
$4.08
57k
108
60
glm-5.2[high]
36% ±5%
$2.84
54k
122
61
gemini-3.5-flash[high]
36% ±4%
$3.45
76k
105
62
gpt-5.6-terra[medium]
35% ±3%
$0.47
12k
25
63
kimi-k2.7-code
31% ±1%
$2.82
59k
149
64
claude-sonnet-5[low]
31% ±1%
$2.19
36k
77
65
claude-sonnet-4.6[high]
30% ±4%
$5.52
76k
134
66
gpt-5.5[low]
27% ±2%
$1.20
9k
28
67
gpt-5.6-terra[low]
24% ±1%
$0.34
9k
21
68
gemini-3.1-pro-preview[high]
12% ±1%
$2.14
28k
76
69
gpt-5.6-luna[medium]
11% ±1%
$0.04
8k
24
70
gpt-5.6-luna[low]
2% ±1%
$0.01
3k
12
0%20%40%60%80%

所有模型均基于 mini-swe-agent 运行以保证结果口径一致。mini-swe-agent 是一个极简的开源软件工程智能体框架,去除了复杂的工具封装与脚手架,让排行榜上的分数差距更真实地反映模型本身的能力,而非 agent 工程的优劣。