Bug-fix benchmark · 24 real bugs · 10 languages
Twenty-four real bugs from real open-source repositories — Go, Rust, C, C++, C#, Java, Python, Ruby, JavaScript, TypeScript. Every comparison here holds the model constant: Benzi, Claude Code and OpenCode all run Claude Sonnet 5; Benzi and DeepSeek's own harness both run DeepSeek v4-flash. Same issue text, same fresh checkout, same test suite. The only thing that changes is the harness driving the model.
How that order is settled. Accuracy first, speed only as the tiebreak. Benzi, Claude Code and DeepSeek Harness each fixed 23 of the bugs they attempted; OpenCode fixed 6. Among the three that tie on accuracy, Benzi is 41% faster than Claude Code on the identical model, and Claude Code is in turn 16% faster than DeepSeek Harness in 24% fewer model calls — that last pair is the one ranking on this page where the two arms are not on the same model, so read it as harness-plus-model, not as a clean harness result.
Check the working. Every task's verbatim prompt, the upstream repository and the real commit that fixed each bug, and all 578 runs ever recorded — failures, superseded runs and the bugs we lose included — are on the full record.
It solved 6 of 24. On thirteen more it was still going when we stopped it, having already spent between 2.4× and 14.2× what Benzi needed to finish the same bug.
Those numbers are its best showing — measured only on the six bugs it actually fixed. Four more it finished with the wrong answer. Thirteen it never finished at all. Nothing here is an average dragged down by the failures; they are excluded entirely.
| bug | Benzi | Claude Code | OpenCode | vs Benzi | vs Claude Code |
|---|---|---|---|---|---|
| mux | 44s / 5t | 61s / 8t | 346s / 51t | 7.8× slower | 5.7× slower |
| commons-cli | 84s / 14t | 67s / 9t | 137s / 24t | 1.6× slower | 2.0× slower |
| cJSON | 87s / 12t | 75s / 17t | 216s / 18t | 2.5× slower | 2.9× slower |
| dayjs | 91s / 11t | 164s / 17t | 222s / 31t | 2.5× slower | 1.4× slower |
| marked | 272s / 49t | 814s / 47t | 919s / 123t | 3.4× slower | 1.1× slower |
| scrapy | 358s / 31t | 400s / 40t | 685s / 40t | 1.9× slower | 1.7× slower |
| total | 937s / 122t | 1582s / 138t | 2526s / 287t | 2.7× slower | 1.6× slower |
Every harness opens more source as bugs get harder. The question is the slope. Each point is one bug; the 24 are laid out easiest to hardest, left to right.
Difficulty is Claude Code's turn count on that bug — a third-party yardstick, so no harness sets its own position on the axis. Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The figure beside each line is its slope: how many extra lines that harness opens per step of difficulty. Hover any point for the bug and its count.
| bug | Benzi Sonnet |
Benzi DeepSeek |
Claude Code Sonnet |
DeepSeek Harness DeepSeek |
|---|---|---|---|---|
| mux | 64 | 64 | 383 | 1,372 |
| commons-cli | 39 | 235 | 456 | 771 |
| addressable | 180 | 240 | 200 | 1,603 |
| jsoup | 61 | 170 | 60 | 616 |
| yaml-cpp | 120 | 383 | 194 | 461 |
| cJSON | 17 | 187 | 105 | 854 |
| dayjs | 64 | 336 | 307 | 610 |
| gson | 78 | 158 | 80 | 516 |
| CsvHelper | 42 | 390 | 335 | 805 |
| semver | 353 | 468 | 563 | 1,857 |
| hashie | 153 | 172 | 300 | 458 |
| money | 681 | 572 | 661 | 1,892 |
| rich | 255 | 514 | 736 | 1,421 |
| fmt | 604 | 315 | 264 | 1,097 |
| sqlglot | 227 | 115 | 800 | 777 |
| scrapy | 386 | 1,580 | 1,259 | 2,127 |
| marked | 905 | 807 | 2,167 | 4,091 |
| http-parser | 449 | 1,201 | 1,270 | 2,635 |
| zod | 841 | 1,476 | 1,956 | 3,608 |
| quartznet | 488 | 1,707 | 1,897 | 4,197 |
| sqlparser | 620 | 1,007 | 1,225 | 2,334 |
| nlohmann/json | 379 | 929 | 1,764 | 5,201 |
| ts-pattern | 1,130 | 1,232 | 1,139 | 1,910 |
| nats-server | 989 | 2,149 | 2,583 | 2,385 |
| all 24 | 9,125 | 16,407 | 20,704 | 43,598 |
Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The green figure in each row is the lowest of the four.
The same 24 bugs in the same order, with wall clock in place of lines read.
Wall clock is raw here — unlike the tables above, Benzi's per-repo index build is not subtracted, so these seconds run slightly higher than the warm figures quoted elsewhere on this page. Each point is that harness's most recent solved run for that bug; unsolved and unfinished runs are left out rather than plotted as fast. Benzi on Sonnet never solved http-parser and the DeepSeek harness never ran nats-server, so those two points are absent and neither enters its fit.
And the same again with dollars on the vertical axis.
Priced at the published per-token rates, same run selection as the chart above it. The two DeepSeek series run an order of magnitude cheaper than the two Sonnet ones, so at this scale they sit close to the baseline — the per-bug figures behind them are in the DeepSeek table further down. What the axis does show is the slope: Claude Code's cost climbs with difficulty faster than any other series here.
| bug | Benzi Sonnet |
Benzi DeepSeek |
Claude Code Sonnet |
DeepSeek Harness DeepSeek |
|---|---|---|---|---|
| mux | $0.39 | $0.036 | $0.25 | $0.023 |
| commons-cli | $0.26 | $0.030 | $0.29 | $0.014 |
| addressable | $0.57 | $0.043 | $0.35 | $0.042 |
| jsoup | $0.23 | $0.025 | $0.30 | $0.051 |
| yaml-cpp | $0.25 | $0.068 | $0.37 | $0.024 |
| cJSON | $0.27 | $0.036 | $0.44 | $0.073 |
| dayjs | $0.21 | $0.038 | $0.55 | $0.040 |
| gson | $0.19 | $0.015 | $0.45 | $0.022 |
| CsvHelper | $0.25 | $0.038 | $0.58 | $0.042 |
| semver | $0.62 | $0.10 | $1.43 | $0.097 |
| hashie | $0.69 | $0.051 | $0.92 | $0.043 |
| money | $0.79 | $0.036 | $0.95 | $0.052 |
| rich | $0.33 | $0.11 | $1.30 | $0.053 |
| fmt | $0.54 | $0.14 | $1.12 | $0.071 |
| sqlglot | $0.62 | $0.033 | $1.23 | $0.059 |
| scrapy | $0.73 | $0.10 | $1.35 | $0.19 |
| marked | $1.08 | $0.20 | $3.59 | $0.13 |
| http-parser | — | $0.36 | $3.44 | $0.23 |
| zod | $1.53 | $0.13 | $2.47 | $0.33 |
| quartznet | $0.95 | $0.23 | $3.99 | $0.31 |
| sqlparser | $0.72 | $0.14 | $3.22 | $0.33 |
| nlohmann/json | $1.03 | $0.17 | $3.68 | $0.44 |
| ts-pattern | $3.74 | $0.31 | $4.33 | $0.053 |
| nats-server | $1.99 | $0.23 | $2.94 | — |
| all 24 | $17.96 | $2.66 | $39.54 | $2.70 |
Priced at published per-token rates. The green figure in each row is the lowest of the four; the two DeepSeek columns are cheaper largely because that model costs roughly twenty times less per token. Blank cells are the two runs that never produced a fix.
This is the real comparison — two harnesses that both finish the job, on the same model. Benzi solves the same bugs in less time, for less money, having read less than half as much of the codebase.
Reading less is the mechanism, not a side effect. Claude Code opens 99 lines per read and takes big grabs; Benzi opens 43, because Benzi's Code Intelligence already knows which lines matter. Across the 22 timed bugs that is 16,851 lines against 7,525 — from almost exactly the same number of reads (171 against 174). Benzi does not look in fewer places. It looks in the same number of places and opens less of each one.
Wall-clock and model calls per bug. Green means Benzi was faster.
| bug | Benzi | Claude Code | OpenCode | Benzi was | lines read CC → Benzi |
|---|---|---|---|---|---|
| muxgo | 44s / 5t | 61s / 8t | 346s / 51t | 27% faster | 383 → 64 |
| addressableruby | 66s / 13t | 56s / 10t | wrong answer | 18% slower | 200 → 180 |
| yaml-cppc++ | 69s / 15t | 92s / 11t | unfinished | 24% faster | 194 → 120 |
| hashieruby | 82s / 15t | 200s / 27t | unfinished | 59% faster | 300 → 153 |
| moneyruby | 84s / 24t | 140s / 30t | wrong answer | 40% faster | 661 → 681 |
| commons-clijava | 84s / 14t | 67s / 9t | 137s / 24t | 25% slower | 456 → 39 |
| cJSONc | 87s / 12t | 75s / 17t | 216s / 18t | 16% slower | 105 → 17 |
| dayjsjs | 91s / 11t | 164s / 17t | 222s / 31t | 45% faster | 307 → 64 |
| gsonjava | 93s / 6t | 134s / 18t | unfinished | 30% faster | 80 → 38 |
| CsvHelperc# | 95s / 15t | 105s / 20t | unfinished | 9% faster | 335 → 42 |
| jsoupjava | 99s / 14t | 80s / 10t | unfinished | 23% slower | 60 → 50 |
| richpython | 115s / 20t | 302s / 36t | unfinished | 62% faster | 736 → 255 |
| semverrust | 138s / 19t | 441s / 23t | unfinished | 69% faster | 563 → 353 |
| nlohmann/jsonc++ | 185s / 33t | 504s / 84t | unfinished | 63% faster | 1,764 → 379 |
| fmtc++ | 216s / 30t | 199s / 37t | unfinished | 9% slower | 264 → 604 |
| markedjs | 272s / 49t | 814s / 47t | 919s / 123t | 67% faster | 2,167 → 905 |
| zodts | 292s / 53t | 438s / 58t | wrong answer | 33% faster | 1,956 → 731 |
| quartznetc# | 321s / 29t | 560s / 61t | unfinished | 43% faster | 1,897 → 488 |
| scrapypython | 358s / 31t | 400s / 40t | 685s / 40t | 10% faster | 1,259 → 385 |
| sqlparserrust | 376s / 33t | 1076s / 61t | unfinished | 65% faster | 1,225 → 620 |
| sqlglotpython | 479s / 23t | 310s / 39t | unfinished | 54% slower | 800 → 227 |
| ts-patternts | 647s / 83t | 1060s / 84t | wrong answer | 39% faster | 1,139 → 1,130 |
| http-parserc | not solved | solved | unfinished | — | — |
| nats-servergo | solved | not solved | not run | — | — |
| 22 timed bugs | 4294s / 547t | 7279s / 747t | — | 41% faster | 16,851 → 7,525 |
Two bugs carry no timing, for opposite reasons. http-parser resists both harnesses: across every run recorded, Claude Code fixes it 1 time in 4 and Benzi 0 in 4 on this model, so a single result there measures luck rather than speed — it is counted as a loss for Benzi. nats-server is the reverse: Benzi has solved it 4 times in 5 attempts and Claude Code 2 in 4, and on this run Benzi fixed it while Claude Code did not. A stopwatch cannot compare a fix against a run that never produced one, so it is reported as a solve rate and kept out of the average — which costs Benzi the win rather than banking it.
DeepSeek shipped its own agent harness on 13 August 2026. It defaults to deepseek-v4-flash — the same model Benzi runs here — so the model is held constant and the harness is the only thing that differs. Run interleaved, one bug at a time, both arms back to back.
| bug | Benzi | DeepSeek Harness | Benzi was | lines read DSH → Benzi |
|---|---|---|---|---|
| muxgo | 89s / 12t | 75s / 13t | 1.2× slower | 1,372 → 64 |
| moneyruby | 95s / 13t | 158s / 20t | 1.7× faster | 1,892 → 572 |
| gsonjava | 114s / 11t | 99s / 20t | 1.2× slower | 516 → 158 |
| hashieruby | 118s / 17t | 168s / 28t | 1.4× faster | 458 → 172 |
| cJSONc | 119s / 22t | 268s / 41t | 2.3× faster | 854 → 187 |
| addressableruby | 125s / 11t | 162s / 25t | 1.3× faster | 1,603 → 240 |
| commons-clijava | 140s / 22t | 65s / 14t | 2.1× slower | 771 → 235 |
| CsvHelperc# | 169s / 18t | 207s / 25t | 1.2× faster | 805 → 390 |
| yaml-cppc++ | 212s / 36t | 154s / 16t | 1.4× slower | 461 → 383 |
| jsoupjava | 232s / 16t | 273s / 34t | 1.2× faster | 616 → 170 |
| dayjsjs | 233s / 16t | 153s / 24t | 1.5× slower | 610 → 336 |
| sqlglotpython | 268s / 21t | 265s / 37t | 1% slower | 777 → 115 |
| richpython | 313s / 62t | 187s / 40t | 1.7× slower | 1,421 → 514 |
| scrapypython | 428s / 33t | 689s / 84t | 1.6× faster | 2,127 → 1,580 |
| semverrust | 436s / 19t | 322s / 33t | 1.4× slower | 1,857 → 468 |
| fmtc++ | 473s / 50t | 335s / 54t | 1.4× slower | 1,097 → 315 |
| nlohmann/jsonc++ | 494s / 52t | 1659s / 124t | 3.4× faster | 5,201 → 929 |
| markedjs | 533s / 74t | 305s / 38t | 1.7× slower | 4,091 → 807 |
| quartznetc# | 551s / 61t | 656s / 90t | 1.2× faster | 4,197 → 1,707 |
| zodts | 583s / 49t | 715s / 98t | 1.2× faster | 3,608 → 1,476 |
| ts-patternts | 767s / 70t | 225s / 26t | 3.4× slower | 1,910 → 1,232 |
| http-parserc | 1000s / 65t | 598s / 79t | 1.7× slower | 2,635 → 1,201 |
| sqlparserrust | 1471s / 54t | 1731s / 90t | 1.2× faster | 2,334 → 1,007 |
| 23 timed bugs | 8964s / 804t | 9470s / 1053t | 6% faster | 41,213 → 14,258 |
This is the closest race on the page, and the honest summary is a tie on the thing that matters: both harnesses fixed all 23. Six percent on wall-clock is inside the noise — Benzi wins 11 bugs and loses 12. What is not close is how much source each opened to get there: 41,213 lines against 14,258, and DeepSeek Harness read more on every single bug. A 24th bug, nats-server, is excluded because neither arm produced a result — Benzi failed it, and the harness run died in cleanup before it could be graded. Benzi's wall-clock here is warm; the per-run index build is subtracted, though at 168s across 23 bugs it barely moves the number (cold, Benzi is still 4% ahead). DeepSeek Harness was v0.1.0-rc.6, released the day these ran.
On DeepSeek v4-flash — a model roughly twenty times cheaper — OpenCode holds up far better than it does on Sonnet: it finishes nine of these ten instead of six of twenty-four. It still ships a fix that fails the tests, and it is still slower, still more model calls, still more expensive, still reads more of the codebase to get there.
| bug | Benzi | OpenCode | Benzi is | lines read OpenCode → Benzi |
|---|---|---|---|---|
| mux | 57s / 10t | 323s / 60t | 5.7× faster | 3407 → 109 |
| addressable | 72s / 6t | 175s / 42t | 2.4× faster | 1696 → 166 |
| yaml-cpp | 118s / 24t | 295s / 33t | 2.5× faster | 1751 → 190 |
| jsoup | 140s / 14t | 634s / 127t | 4.5× faster | 1049 → 163 |
| commons-cli | 76s / 15t | 359s / 56t | 4.7× faster | 2563 → 23 |
| money | 157s / 16t | 93s / 29t — wrong | solved it | 652 → 222 |
| CsvHelper | 101s / 16t | 245s / 41t | 2.4× faster | 701 → 355 |
| cJSON | 125s / 23t | 179s / 43t | 1.4× faster | 1137 → 570 |
| gson | 143s / 15t | 160s / 26t | 1.1× faster | 1137 → 355 |
| dayjs | 267s / 22t | 664s / 43t | 2.5× faster | 1176 → 331 |
| the 9 both solved | 1099s / 145t | 3033s / 471t | 2.8× faster | 14,617 → 2,262 |
The reading gap is even wider here: across the nine both harnesses solved, OpenCode pulled 14,617 lines of source and Benzi 2,262. Same model, same bugs — one of them reads the codebase to find the answer, the other looks it up.
Totals cover the nine bugs OpenCode actually fixed. On the tenth, money, it stopped after 93 seconds with the bug still there — averaging that in would credit it for a fast give-up and understate Benzi (2.5× instead of 2.8×). The same rule removes nats-server from the Sonnet table above, where Claude Code failed and Benzi did not, and it is why http-parser counts as a loss for Benzi: a run that never fixed the bug cannot tell you how long fixing it takes.
Why that counts as a loss. money is graded against OpenCode even though it was the quickest run in the row, because the suite still fails afterwards and this benchmark settles order on correctness first, with speed only as the tiebreak. The same ordering is why OpenCode's slow runs are not left to run out: once one passes twice the slower of the other two arms on the same bug it is stopped and recorded as unfinished — not a failure, but not a completion either, and kept out of every average. A harness that is behind on accuracy and on time, model calls and cost at once has already answered the question the run was asking.
The exact text handed to every agent, unedited.