29 Sep 2026 · 5 min read

What Concourse Bench v1 reveals about AI buyer error rates

The debate about AI media buyers has largely been conducted without data. Vendors make claims about capability, clients weigh those claims against scepticism, and the industry moves forward on a combination of early pilots and competitive pressure. What has been missing is a rigorous, independent evaluation of how AI buying systems actually perform across standardised tasks. That is what Concourse Bench v1 provides.

What did Concourse Bench v1 test?

The benchmark evaluated eight model-harness combinations across a structured set of media buying scenarios. The combinations were named Sol & Fable, Opus, Astra & Sonnet, Terra, and Luna & Haiku. Each was tested against the same buying tasks, with performance measured on whether the attempt reached a completed deal that met the brief, and whether the deal terms hit minimum quality standards.

The methodology evaluates the complete system, not the model in isolation. This is significant because it reflects how AI buyers are actually deployed: as model-harness combinations in which the software layer around the model shapes what the model can do and what the market sees.

What was the overall completion rate?

Of 96 registered attempts across all eight combinations, 55 reached completion. Five attempts were stopped by technical issues before assessment could take place, leaving 91 assessable attempts of which 55 produced a completed buy. That is a completion rate of roughly 60% across the field.

This aggregate figure masks significant variance between combinations. Sol & Fable completed 12 of 12 attempts: a 100% completion rate. Opus completed 10 of 12. Astra & Sonnet completed 9 of 12. Terra completed 3 of 12. Luna & Haiku completed 0 of 12. The range from complete reliability to complete failure across the field, using the same tasks and the same assessment criteria, is the central finding of the benchmark.

What were the failure modes?

The benchmark identified three categories of failure. Budget overrun: the buyer committed to terms that exceeded the available budget, either because it failed to account for existing commitments or because the harness did not enforce budget limits as a hard constraint. Missed brief requirement: the buyer committed to terms that did not meet a specified requirement in the brief, such as a targeting parameter, a placement restriction, or a performance floor. Technical interruption: the harness encountered an error during the negotiation sequence and could not recover, causing the attempt to terminate without a completed deal.

These failure modes are not random. Budget overrun and missed brief requirement are governance failures: the harness did not enforce the constraints it was supposed to enforce. Technical interruption is a reliability failure: the harness was not robust enough to handle the conditions it encountered. The distribution of failures across combinations reflects the underlying quality of harness design, not just model capability differences.

What does the viewability finding mean?

The benchmark examined the deal terms in 110 contracts committed by AI buyers across the test scenarios. Of those 110 contracts, 95 had viewability floors below 70%. This finding cuts across all combinations and is independent of completion rate. Even buyers that successfully completed their negotiation attempts frequently committed to inventory at viewability floors that most sophisticated advertisers would consider insufficient.

This is a quality finding, not just a completion finding. A buyer can successfully complete a negotiation, in the sense of reaching a committed deal within budget, while simultaneously committing to inventory terms that fail a basic quality standard. The Concourse Bench methodology captures both dimensions. The viewability finding suggests that in most of the tested scenarios, viewability was not enforced as a minimum contract standard in the buyer's harness.

What do the API costs reveal?

The benchmark recorded API costs for each completed buy, ranging from $0.74 to $14.38 per completed transaction. This range reflects the significant variation in how computationally intensive different model-harness combinations are when executing a media buy. The cost of running a completed buy is a material operational consideration for any organisation scaling AI buying across a large campaign portfolio.

The cost data also provides an indirect signal about processing approach. A buyer that spends significantly more compute per transaction may be doing more reasoning steps per negotiation round, which could explain higher completion rates in complex scenarios. A buyer that spends less may be taking simpler approaches that break down in more demanding conditions. The relationship between cost and quality is not linear, and the benchmark data should be read with that complexity in mind.

What should practitioners take from the Concourse Bench v1 results?

Three conclusions are well-supported by the data. First, completion rates vary enormously across model-harness combinations. An organisation that deploys an AI buyer without testing it against standardised tasks is taking a significant performance risk. The range between 12/12 and 0/12 across the same task set is not a rounding error.

Second, completing a deal is not the same as completing a quality deal. The viewability finding demonstrates that buyers can reach committed deals that fail basic quality standards. Evaluation criteria need to include quality compliance, not just completion rate.

Third, the harness is the variable that explains most of the performance difference. Similar model tiers in different harness configurations produced different outcomes. Practitioners evaluating AI buying systems should treat harness design as a primary evaluation criterion, not a secondary one. The Concourse Bench methodology is designed to surface exactly these distinctions, and the v1 results are the first rigorous public evidence base for how consequential those distinctions are in practice.

All articles

Speak to the team