Benchmark 01R · updated October 7, 2026
Can AI do a quantity takeoff?
Reading a drawing is one thing. Producing a number someone can price is another. We handed frontier models a complete drawing set and the estimator's own bill of quantities with every quantity stripped out, and asked them to fill it back in — then scored it line by line against a takeoff measured by a chartered civil engineer.
Not one of them produced a bill you could price. The best raw model was right on 48% of the lines.
The same models inside a harness
One session per set per harness unless the row says otherwise, with the drawings and the blank bill in a folder the agent can open and work through however it likes. Both sets, 130 lines, on the brief that still lets a session leave a line blank. The last column compares each row with the same model through the API on that same brief, where it sat raw. Click a row for the split by set. Not ranked yet: Cowork · Claude Opus 5 (66.7% on electrical only) and Cursor · GPT 5.6 Sol (50.0% on electrical only) — they have sat one set so far and join the board when the second is in.
| # | Harness | Model | Correct | Lines | Declined | Counts | Measured | Cost | vs raw | Details |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Cowork electrical + plumbing · both sets in one session | Claude Opus 5.5 | 53.1% | 130 | 3 | 69%of 70 | 30%of 50 | subscription | +24.6 pts API, same brief: 28.5% | |
Cowork · Claude Opus 5.5 by set69 of 130 lines inside tolerance, 3 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 2 | Cowork electrical + plumbing · both sets in one session | Claude Sonnet 5.5 | 49.2% | 130 | 3 | 69%of 70 | 20%of 50 | subscription | +45.4 pts API, same brief: 3.9% | |
Cowork · Claude Sonnet 5.5 by set64 of 130 lines inside tolerance, 3 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 3= | ChatGPT Work electrical + plumbing · both sets in one session · self-built tooling | GPT 6 Astra | 46.2% | 130 | 8 | 61%of 70 | 28%of 47 | subscription | +7.7 pts API, same brief: 38.5% | |
ChatGPT Work · GPT 6 Astra by set60 of 130 lines inside tolerance, 8 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 3= | ChatGPT Work electrical + plumbing · both sets in one session · self-built tooling | GPT 6.1 Sol | 46.2% | 130 | 16 | 63%of 68 | 30%of 40 | subscription | +9.2 pts API, same brief: 36.9% | |
ChatGPT Work · GPT 6.1 Sol by set60 of 130 lines inside tolerance, 16 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 5 | Cowork electrical + plumbing · both sets in one session | Claude Fable 5.1 | 43.9% | 130 | 5 | 67%of 70 | 8%of 48 | subscription | no raw run on both sets | |
Cowork · Claude Fable 5.1 by set57 of 130 lines inside tolerance, 5 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 6 | Claude Code electrical + plumbing · 1 session each | Claude Opus 5 | 43.1% | 130 | 2 | 67%of 70 | 6%of 51 | subscription | +12.7 pts API, same brief: 30.4% | |
Claude Code · Claude Opus 5 by set56 of 130 lines inside tolerance, 2 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 7 | Cursor electrical + plumbing · 1 session each | Claude Opus 5 | 40.8% | 130 | 24 | 66%of 68 | 6%of 31 | subscription | +10.4 pts API, same brief: 30.4% | |
Cursor · Claude Opus 5 by set53 of 130 lines inside tolerance, 24 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 8 | ChatGPT Work electrical + plumbing · 1 session each | GPT 5.6 Sol | 35.4% | 130 | — | 55%of 71 | 4%of 52 | subscription | +0.4 pts API, same brief: 35.0% | |
ChatGPT Work · GPT 5.6 Sol by set46 of 130 lines inside tolerance. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 9 | Codex electrical + plumbing · 1 session each | GPT 5.6 Sol | 32.3% | 130 | 2 | 51%of 71 | 2%of 50 | subscription | −2.7 pts API, same brief: 35.0% | |
Codex · GPT 5.6 Sol by set42 of 130 lines inside tolerance, 2 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 10= | ChatGPT Work electrical + plumbing · both sets in one session | GPT 6 Sol | 30.8% | 130 | 62 | 59%of 58 | 33%of 3 | subscription | +3.1 pts API, same brief: 27.7% | |
ChatGPT Work · GPT 6 Sol by set40 of 130 lines inside tolerance, 62 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 10= | Cursor electrical + plumbing · 1 session each | Grok 4.6 | 30.8% | 130 | 14 | 50%of 69 | 0%of 41 | subscription | no raw run on both sets | |
Cursor · Grok 4.6 by set40 of 130 lines inside tolerance, 14 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 12 | Muse electrical + plumbing · both sets in one session | Muse Spark 1.3 | 16.2% | 130 | 93 | 57%of 30 | 0%of 1 | subscription | no raw run on both sets | |
Muse · Muse Spark 1.3 by set21 of 130 lines inside tolerance, 93 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 13 | ChatGPT Work electrical + plumbing · 1 session each | GPT 5.6 Terra | 13.1% | 130 | 96 | 50%of 34 | n/a | subscription | −8.5 pts API, same brief: 21.5% | |
ChatGPT Work · GPT 5.6 Terra by set17 of 130 lines inside tolerance, 96 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
Raw models
No harness. The model, the drawings and the blank bill, nothing to open and nowhere to work. Every line must carry a number: where the drawings don't support a firm figure, the model gives its best estimate and flags it, the way an estimator would. Scored over every line, across both sets. Click a row for the split. Not shown: Claude Fable 5.1, Grok 4.6, Grok 4.5 and Gemini 3.8 Flash, which could sit only one of the two sets, so they are withheld rather than ranked on fewer lines.
| # | Model | Lab | Correct | Lines | Flagged | Counts | Measured | Cost | Details |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT 6 Astra openai/gpt-6-astra | OpenAI | 48.5% | 130 | 70 | 66%of 71 | 19%of 52 | $2.95 | |
GPT 6 Astra by set63 of 130 lines inside tolerance, 70 flagged as estimates. Allowed to leave lines blank, it scored 38.5% and declined 51. $2.95 to fill the bill once.
| |||||||||
| 2 | GPT 6.1 Sol openai/gpt-6.1-sol | OpenAI | 40.0% | 130 | 61 | 58%of 71 | 10%of 52 | $0.54 | |
GPT 6.1 Sol by set52 of 130 lines inside tolerance, 61 flagged as estimates. Allowed to leave lines blank, it scored 36.9% and declined 38. $0.54 to fill the bill once.
| |||||||||
| 3 | Claude Opus 5.5 anthropic/claude-opus-5.5 | Anthropic | 36.9% | 130 | 98 | 51%of 71 | 13%of 52 | $0.65 | |
Claude Opus 5.5 by set48 of 130 lines inside tolerance, 98 flagged as estimates. Allowed to leave lines blank, it scored 28.5% and declined 61. $0.65 to fill the bill once.
| |||||||||
| 4 | Claude Opus 4.7 anthropic/claude-opus-4.7 | Anthropic | 33.9% | 130 | 93 | 51%of 71 | 6%of 52 | $0.73 | |
Claude Opus 4.7 by set44 of 130 lines inside tolerance, 93 flagged as estimates. Allowed to leave lines blank, it scored 20.0% and declined 88. $0.73 to fill the bill once.
| |||||||||
| 5 | GPT 5.6 Sol openai/gpt-5.6-sol | OpenAI | 32.7% | 130 | 63 | 50%of 71 | 2%of 52 | $0.55 | |
GPT 5.6 Sol by set43 of 130 lines inside tolerance, 63 flagged as estimates. Allowed to leave lines blank, it scored 35.0% and declined 3. $0.55 to fill the bill once.
| |||||||||
| 6 | Claude Opus 5 anthropic/claude-opus-5 | Anthropic | 32.3% | 130 | 94 | 47%of 71 | 7%of 52 | $0.91 | |
Claude Opus 5 by set42 of 130 lines inside tolerance, 94 flagged as estimates. Allowed to leave lines blank, it scored 30.4% and declined 8. $0.91 to fill the bill once.
| |||||||||
| 7 | GPT 5.6 Terra openai/gpt-5.6-terra | OpenAI | 31.5% | 130 | 80 | 46%of 71 | 6%of 52 | $0.51 | |
GPT 5.6 Terra by set41 of 130 lines inside tolerance, 80 flagged as estimates. Allowed to leave lines blank, it scored 21.5% and declined 90. $0.51 to fill the bill once.
| |||||||||
| 8 | Gemini 3.6 Flash google/gemini-3.6-flash | 31.1% | 130 | — | 46%of 71 | 6%of 52 | $0.11 | ||
Gemini 3.6 Flash by set41 of 130 lines inside tolerance. Allowed to leave lines blank, it scored 27.7% and declined 0. $0.11 to fill the bill once.
| |||||||||
| 9 | Claude Fable 5 anthropic/claude-fable-5 | Anthropic | 30.8% | 130 | 93 | 45%of 71 | 6%of 52 | $1.93 | |
Claude Fable 5 by set40 of 130 lines inside tolerance, 93 flagged as estimates. Allowed to leave lines blank, it scored 18.5% and declined 96. $1.93 to fill the bill once.
| |||||||||
| 10 | Claude Opus 4.8 anthropic/claude-opus-4.8 | Anthropic | 30.0% | 130 | 107 | 46%of 71 | 2%of 52 | $0.85 | |
Claude Opus 4.8 by set39 of 130 lines inside tolerance, 107 flagged as estimates. Allowed to leave lines blank, it scored 20.8% and declined 96. $0.85 to fill the bill once.
| |||||||||
| 11 | Claude Haiku 4.5 anthropic/claude-haiku-4.5 | Anthropic | 29.2% | 130 | 89 | 46%of 71 | 0%of 52 | $0.13 | |
Claude Haiku 4.5 by set38 of 130 lines inside tolerance, 89 flagged as estimates. Allowed to leave lines blank, it scored 13.9% and declined 94. $0.13 to fill the bill once.
| |||||||||
| 12= | Claude Sonnet 5 anthropic/claude-sonnet-5 | Anthropic | 28.5% | 130 | 98 | 45%of 71 | 0%of 52 | $0.49 | |
Claude Sonnet 5 by set37 of 130 lines inside tolerance, 98 flagged as estimates. Allowed to leave lines blank, it scored 11.5% and declined 112. $0.49 to fill the bill once.
| |||||||||
| 12= | GPT 5.4 Mini openai/gpt-5.4-mini | OpenAI | 28.5% | 130 | 121 | 42%of 71 | 6%of 52 | $0.15 | |
GPT 5.4 Mini by set37 of 130 lines inside tolerance, 121 flagged as estimates. Allowed to leave lines blank, it scored 15.4% and declined 103. $0.15 to fill the bill once.
| |||||||||
| 12= | GPT 6 Sol openai/gpt-6-sol | OpenAI | 28.5% | 130 | 46 | 44%of 71 | 2%of 52 | $0.51 | |
GPT 6 Sol by set37 of 130 lines inside tolerance, 46 flagged as estimates. Allowed to leave lines blank, it scored 27.7% and declined 67. $0.51 to fill the bill once.
| |||||||||
| 15= | Gemini 3.1 Flash Lite google/gemini-3.1-flash-lite | 27.7% | 130 | 1 | 45%of 71 | 0%of 52 | $0.02 | ||
Gemini 3.1 Flash Lite by set36 of 130 lines inside tolerance, 1 flagged as estimates. Allowed to leave lines blank, it scored 23.1% and declined 53. $0.02 to fill the bill once.
| |||||||||
| 15= | Gemini 3.7 Flash google/gemini-3.7-flash | 27.7% | 130 | — | 44%of 71 | 0%of 52 | $0.08 | ||
Gemini 3.7 Flash by set36 of 130 lines inside tolerance. Allowed to leave lines blank, it scored 27.7% and declined 0. $0.08 to fill the bill once.
| |||||||||
| 17= | Claude Sonnet 5.5 anthropic/claude-sonnet-5.5 | Anthropic | 26.9% | 130 | 105 | 44%of 71 | 2%of 52 | $0.29 | |
Claude Sonnet 5.5 by set35 of 130 lines inside tolerance, 105 flagged as estimates. Allowed to leave lines blank, it scored 3.9% and declined 119. $0.29 to fill the bill once.
| |||||||||
| 17= | Gemini 2.5 Pro google/gemini-2.5-pro | 26.9% | 130 | 13 | 39%of 71 | 4%of 52 | $0.29 | ||
Gemini 2.5 Pro by set35 of 130 lines inside tolerance, 13 flagged as estimates. Allowed to leave lines blank, it scored 26.2% and declined 17. $0.29 to fill the bill once.
| |||||||||
| 17= | Gemini 3.1 Pro Preview google/gemini-3.1-pro-preview | 26.9% | 130 | 130 | 43%of 71 | 3%of 52 | $0.22 | ||
Gemini 3.1 Pro Preview by set35 of 130 lines inside tolerance, 130 flagged as estimates. Allowed to leave lines blank, it scored 10.8% and declined 107. $0.22 to fill the bill once.
| |||||||||
| 20 | GPT 5.5 openai/gpt-5.5 | OpenAI | 26.5% | 130 | 55 | 39%of 71 | 4%of 52 | $0.63 | |
GPT 5.5 by set35 of 130 lines inside tolerance, 55 flagged as estimates. Allowed to leave lines blank, it scored 26.2% and declined 46. $0.63 to fill the bill once.
| |||||||||
| 21 | Mistral Large 2512 mistralai/mistral-large-2512 | Mistral | 26.2% | 130 | 19 | 39%of 71 | 4%of 52 | $0.03 | |
Mistral Large 2512 by set34 of 130 lines inside tolerance, 19 flagged as estimates. Allowed to leave lines blank, it scored 19.2% and declined 70. $0.03 to fill the bill once.
| |||||||||
| 22 | Gemini 2.5 Flash google/gemini-2.5-flash | 25.4% | 130 | 46 | 41%of 71 | 2%of 52 | $0.07 | ||
Gemini 2.5 Flash by set33 of 130 lines inside tolerance, 46 flagged as estimates. Allowed to leave lines blank, it scored 23.8% and declined 21. $0.07 to fill the bill once.
| |||||||||
| 23= | Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 24.6% | 130 | 21 | 38%of 71 | 4%of 52 | $0.36 | |
Claude Sonnet 4.5 by set32 of 130 lines inside tolerance, 21 flagged as estimates. Allowed to leave lines blank, it scored 20.0% and declined 61. $0.36 to fill the bill once.
| |||||||||
| 23= | GPT 5.6 Luna openai/gpt-5.6-luna | OpenAI | 24.6% | 130 | 71 | 37%of 71 | 2%of 52 | $0.05 | |
GPT 5.6 Luna by set32 of 130 lines inside tolerance, 71 flagged as estimates. Allowed to leave lines blank, it scored 24.6% and declined 78. $0.05 to fill the bill once.
| |||||||||
| 25 | GPT 4.1 openai/gpt-4.1 | OpenAI | 22.3% | 130 | 22 | 46%of 54 | 0%of 6 | $0.14 | |
GPT 4.1 by set29 of 130 lines inside tolerance, 22 flagged as estimates. Allowed to leave lines blank, it scored 24.6% and declined 55. $0.14 to fill the bill once.
| |||||||||
| 26 | Gemini 2.5 Flash Lite google/gemini-2.5-flash-lite | 16.9% | 130 | 53 | 40%of 47 | 6%of 47 | $0.01 | ||
Gemini 2.5 Flash Lite by set22 of 130 lines inside tolerance, 53 flagged as estimates. Allowed to leave lines blank, it scored 30.0% and declined 11. $0.01 to fill the bill once.
| |||||||||
| 27 | GPT 4o openai/gpt-4o | OpenAI | 3.1% | 130 | 16 | 9%of 35 | 0%of 23 | $0.15 | |
GPT 4o by set4 of 130 lines inside tolerance, 16 flagged as estimates. Allowed to leave lines blank, it scored 14.6% and declined 57. $0.15 to fill the bill once.
| |||||||||
| 28 | Mistral Medium 3.1 mistralai/mistral-medium-3.1 | Mistral | 2.3% | 130 | 127 | 0%of 71 | 0%of 52 | $0.02 | |
Mistral Medium 3.1 by set3 of 130 lines inside tolerance, 127 flagged as estimates. Allowed to leave lines blank, it scored 0.0% and declined 130. $0.02 to fill the bill once.
| |||||||||
| 29 | GPT 5 Nano openai/gpt-5-nano | OpenAI | 0.0% | 130 | 130 | 0%of 71 | 0%of 52 | $0.01 | |
GPT 5 Nano by set0 of 130 lines inside tolerance, 130 flagged as estimates. Allowed to leave lines blank, it scored 0.0% and declined 130. $0.01 to fill the bill once.
| |||||||||
Reading the columns
Flagged
Lines the model priced but marked as an estimate rather than a figure it stands behind. On the raw board every line must carry a number, so this is where a model's doubt shows. A flagged line is scored like any other, and the count sits beside the score so you can see how much of the bill the model itself would want checked.
The in-app board still lets a session leave a line blank, so its Declined column counts the lines it did not price. A declined line earns nothing.
Counts
Lines you get by counting tagged items on a plan — fixtures, fittings, outlets. Scored exactly: the number matches the estimator's or it does not. Most of the bill is this, and it is the work most people want to hand over first.
Measured
Lines you get by measuring off the drawing — pipe and cable runs by diameter, each a sum over dozens of branches. Scored inside ±5%, a tolerance we state rather than one anyone measured. This is where the takeoff is actually won, and it is where every model on this page still falls over.
Read the rate with the count beside it. On the raw board every line carries a number, so each model's rates are over the same lines. The in-app board still lets a session decline, and its rates are over the lines it actually priced, so declining the hard ones raises the percentage. Half of eight measured lines attempted is a worse takeoff than three per cent of fifty, and it prints as the better number. The count under each rate is how many lines it is over.
Updated 29 September 2026: the raw board now requires a number on every line. Until today a model could decline a line it could not read, and a declined line scored as a miss. Some models declined most of the bill, which hid how much of it they could actually read. Every raw model was re-run under the new rule, once each, and each row still shows what it scored when it could decline. The in-app board is unchanged.
Updated 27 September 2026: four plumbing rows used only for scope and combined-count checks are excluded from individual accuracy and decline counts. All results on this page have been recalculated consistently. The takeoff keys were checked by the benchmark author; they have not had an independent second estimator pass.
Correct covers every scored quantity line, including the few read off a schedule that neither column above breaks out. Cost is what one pass over the bill cost through the API; an agent session runs on a flat subscription, so there is no per-run figure to print. Every model on this board sat every set, over the same 130 lines; a model that cannot sit one of them is withheld rather than ranked on fewer. Every number comes from the same scorer that runs internally, off run records kept on disk. Drawings are never published. Full method and the raw results: ContractorOS.