CursorBench 3.2

We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better.

More about CursorBench
A scatter and line chart comparing Fable 5, Opus 5, Opus 4.8, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5, Sonnet 5, GLM 5.2, Composer 2.5, Gemini 3.6 Flash, Gemini 3.5 Flash, Kimi K3, and Kimi K2.7 Code scores against average cost per task.75%CursorBench 3.2 score70%65%60%55%50%45%$18$15$12$9$6$3$0Average cost per taskGrok 4.6Opus 5Fable 5Kimi K3GPT-5.6 SolSonnet 5Composer 2.5Gemini 3.6 FlashGPT-5.6 TerraGPT-5.6 Luna
Model
1Grok 4.6 Extra High70.8%$2.8141,13646
2Fable 5 Max70.5%$17.32103,52572
3Opus 5 Max70.0%$8.2361,83878
4Grok 4.6 High69.9%$2.3432,44939
5Opus 5 Extra High69.3%$7.3554,23972
6Fable 5 Extra High68.4%$11.7364,97156
7GPT-5.6 Sol Max67.2%$5.6928,32048
8Grok 4.6 Medium67.1%$1.2817,94229
9Opus 5 High66.7%$3.9127,93248
10Fable 5 High66.5%$8.7743,74748
11Fable 5 Medium65.2%$6.8030,36641
12GPT-5.6 Terra Max64.9%$2.3132,96947
13GPT-5.6 Sol Extra High64.5%$3.8819,69938
14Opus 5 Medium64.3%$3.2923,61244
15GPT-5.6 Sol High63.5%$2.7913,86732
16Opus 5 Low62.8%$2.5518,52937
17Opus 4.8 Max62.3%$5.7771,41144
18Fable 5 Low62.1%$4.4618,18231
19Sonnet 5 Max61.5%$4.3092,88286
20GPT-5.6 Luna Max61.1%$0.3987,97361
21Grok 4.6 Low61.0%$0.7010,65823
22Kimi K3 Max60.8%$2.7038,42857
23GPT-5.6 Sol Medium60.0%$1.959,74727
24Kimi K3 High59.7%$1.8926,84647
25Opus 4.8 Extra High59.4%$4.5051,12140
26GPT-5.6 Terra Extra High59.2%$1.1516,08929
27Sonnet 5 Extra High58.7%$2.7752,87167
28GPT-5.5 High58.4%$2.0512,18328
29GPT-5.5 Extra High58.4%$2.8517,53432
30Opus 4.8 High58.0%$3.1533,54833
31GPT-5.6 Luna Extra High57.7%$0.2322,48048
32Sonnet 5 High56.9%$2.1339,48357
33GPT-5.6 Luna High56.8%$0.1615,14140
34Opus 4.8 Medium56.1%$2.8128,38432
35Composer 2.556.1%$0.4414,28633
36GLM 5.2 Max55.0%$1.7635,94658
37GPT-5.6 Terra High54.2%$0.719,46823
38GPT-5.5 Medium53.8%$1.518,52225
39Gemini 3.6 Flash High53.5%$1.5630,43664
40Opus 4.8 Low53.1%$2.0219,62427
41GPT-5.6 Sol Low52.6%$1.015,10419
42Sonnet 5 Medium52.4%$1.4426,20046
43GLM 5.2 High51.5%$1.1921,82949
44Gemini 3.6 Flash Medium51.2%$1.4828,51162
45Kimi K3 Low50.5%$0.9913,00733
46GPT-5.6 Terra Medium50.3%$0.496,22220
47Kimi K2.7 Code49.7%$1.4331,24758
48Gemini 3.5 Flash48.8%$2.2046,70277
49GPT-5.6 Luna Medium47.7%$0.087,09528
50Sonnet 5 Low47.7%$0.8716,26933
51Gemini 3.6 Flash Low47.4%$1.1320,52950
52GPT-5.6 Terra Low46.9%$0.425,31219
53GPT-5.5 Low46.6%$0.985,16820
54GPT-5.6 Luna Low37.6%$0.033,20917

Changelog

Reporting

  • Updated Sonnet 5 results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Terra and Luna results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Sol, Terra, and Luna results to account for cache write costs.

Tasks

  • CursorBench 3.2
    • Introduced instruction following and advanced tool use problems.

Tasks

  • CursorBench 3.1
    • Introduced problems focused on codebase understanding, bugfinding, planning, and code review.
    • Improved grading criteria for some edit tasks.

Tasks

  • CursorBench 3.0
    • Initial set of tasks focused on edit, refactor, and bugfix problems.

Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.