Benchmark 01R · updated October 7, 2026

Can AI do a quantity takeoff?

Reading a drawing is one thing. Producing a number someone can price is another. We handed frontier models a complete drawing set and the estimator's own bill of quantities with every quantity stripped out, and asked them to fill it back in — then scored it line by line against a takeoff measured by a chartered civil engineer.

Not one of them produced a bill you could price. The best raw model was right on 48% of the lines.

130
Lines to price
29
Models tested
15
Harnesses tested
0
Priceable bills

The same models inside a harness

One session per set per harness unless the row says otherwise, with the drawings and the blank bill in a folder the agent can open and work through however it likes. Both sets, 130 lines, on the brief that still lets a session leave a line blank. The last column compares each row with the same model through the API on that same brief, where it sat raw. Click a row for the split by set. Not ranked yet: Cowork · Claude Opus 5 (66.7% on electrical only) and Cursor · GPT 5.6 Sol (50.0% on electrical only) — they have sat one set so far and join the board when the second is in.

# Harness Model Correct Lines Declined Counts Measured Cost vs raw Details

Raw models

No harness. The model, the drawings and the blank bill, nothing to open and nowhere to work. Every line must carry a number: where the drawings don't support a firm figure, the model gives its best estimate and flags it, the way an estimator would. Scored over every line, across both sets. Click a row for the split. Not shown: Claude Fable 5.1, Grok 4.6, Grok 4.5 and Gemini 3.8 Flash, which could sit only one of the two sets, so they are withheld rather than ranked on fewer lines.

# Model Lab Correct Lines Flagged Counts Measured Cost Details

Reading the columns

Flagged

Lines the model priced but marked as an estimate rather than a figure it stands behind. On the raw board every line must carry a number, so this is where a model's doubt shows. A flagged line is scored like any other, and the count sits beside the score so you can see how much of the bill the model itself would want checked.

The in-app board still lets a session leave a line blank, so its Declined column counts the lines it did not price. A declined line earns nothing.

Counts

Lines you get by counting tagged items on a plan — fixtures, fittings, outlets. Scored exactly: the number matches the estimator's or it does not. Most of the bill is this, and it is the work most people want to hand over first.

Measured

Lines you get by measuring off the drawing — pipe and cable runs by diameter, each a sum over dozens of branches. Scored inside ±5%, a tolerance we state rather than one anyone measured. This is where the takeoff is actually won, and it is where every model on this page still falls over.

Read the rate with the count beside it. On the raw board every line carries a number, so each model's rates are over the same lines. The in-app board still lets a session decline, and its rates are over the lines it actually priced, so declining the hard ones raises the percentage. Half of eight measured lines attempted is a worse takeoff than three per cent of fifty, and it prints as the better number. The count under each rate is how many lines it is over.

Updated 29 September 2026: the raw board now requires a number on every line. Until today a model could decline a line it could not read, and a declined line scored as a miss. Some models declined most of the bill, which hid how much of it they could actually read. Every raw model was re-run under the new rule, once each, and each row still shows what it scored when it could decline. The in-app board is unchanged.

Updated 27 September 2026: four plumbing rows used only for scope and combined-count checks are excluded from individual accuracy and decline counts. All results on this page have been recalculated consistently. The takeoff keys were checked by the benchmark author; they have not had an independent second estimator pass.

Correct covers every scored quantity line, including the few read off a schedule that neither column above breaks out. Cost is what one pass over the bill cost through the API; an agent session runs on a flat subscription, so there is no per-run figure to print. Every model on this board sat every set, over the same 130 lines; a model that cannot sit one of them is withheld rather than ranked on fewer. Every number comes from the same scorer that runs internally, off run records kept on disk. Drawings are never published. Full method and the raw results: ContractorOS.