28 Sep 2026 · 4 min read
Beyond completion: the four measures that make AI media buying evaluable
The first question anyone asks about an AI media buyer is whether it can close a buy. That is the right first question. It is the wrong last question. A buyer that completes 100% of its attempts may still fail to satisfy the brief, accept low viewability floors, produce inconsistent results across attempts, or cost more across a full workload than the headline figure suggests. Completion rate is the threshold. Four other measures determine whether a buyer that clears the threshold is actually deployable.
TL;DR: Concourse Bench v1 (concourse.agency) tested eight AI buyer models across 96 registered attempts at a fictional UK retail brief. Two models completed all 12 attempts: GPT-5.6 Sol and Claude Fable 5.1. Both cleared the completion threshold. The four measures that then differentiate between them, and between any deployable buyer system, are execution (brief compliance, not just deal completion), commercial quality (contract terms, not just spend), reliability (consistent performance across fresh attempts), and measured efficiency (all-attempt cost, not per-completed-buy cost).
Execution: did the buy satisfy the brief, not just the deal?
A deal closes when contracts are signed and a spend commitment is recorded. A brief is satisfied when the contracted packages, adjusted for audience overlap, clear the reach requirement within the budget constraint. These are not the same event. In Concourse v1, one buyer in the firmer-seller condition signed two valid contracts and returned a completion state. Net forecast reach was 124,511 against a brief target of 137,156. The shortfall was 9.2%. The deal closed. The brief was not satisfied.
Execution, reported as a distinct measure from deal completion, would have surfaced this. A buyer system that distinguishes between 'deals closed' and 'brief satisfied' produces a different signal for a campaign planner than one that returns a single completion indicator. In a production deployment, where the campaign window is finite and the budget is committed, that distinction determines whether a shortfall is caught before or after the window closes.
Commercial quality: what terms did the buy commit to?
A completed buy with weak contract terms is a liability, not an asset. In Concourse v1, 95 of 110 contracts produced in 55 completed buys carried viewability floors below 70%. These were completed buys. They closed. The contracts they produced were systematically under-protected on viewability.
Commercial quality as a distinct measure tracks viewability floors, make-good provisions, audience delivery guarantees, and any other contract term that carries risk or value beyond the headline spend commitment. It is the measure that surfaces the gap between a buyer that commits spend efficiently and a buyer that commits spend against terms that protect the campaign. GPT-6 Astra and GPT-5.6 Sol committed identical spend of £18,706.89 against the same brief. Their contract terms were not identical. Spend efficiency does not surface that. Commercial quality does.
Reliability: does the buyer perform consistently, not just occasionally?
AI buyer behaviour is not deterministic. A model that completes a buy on its first attempt is not guaranteed to complete the same buy on a second or third fresh attempt. Concourse v1 ran three fresh attempts per model per buying situation. That structure made reliability measurable. GPT-5.6 Sol and Claude Fable 5.1 both completed all 12 registered attempts across all four situation variants. Claude Opus 5 completed 10 of its 10 assessed attempts. Claude Haiku 4.5 and Claude Luna completed 0 of their assessed attempts.
Peak performance and reliable performance are different deployment considerations. A buyer that completes a reference buy every time but fails across tighter-seller variants may be adequate for standard deployments and unsuitable for constrained inventory environments. Reliability, measured across multiple fresh attempts at multiple situations, is the dimension that determines deployment scope. A single-attempt evaluation cannot produce it.
Measured efficiency: what does the full workload cost?
Cost per completed buy is the figure most frequently cited. It is the most misleading figure for deployment decisions. A buyer that completes 50% of its attempts costs twice as much per successful outcome as its per-completed-buy figure suggests, because the cost of all failed attempts is real infrastructure spend against zero output.
Measured efficiency in Concourse v1 is reported as all-attempt cost: the total compute cost across every registered attempt, divided by the number of models, situations, and attempts in each workload. Sol's all-attempt cost across its full 12-attempt workload was $35.13. Fable's was $129.03. Both models completed all 12 attempts, so their per-completed-buy cost is identical to their all-attempt cost. For models with lower completion rates, the two figures would diverge significantly. The all-attempt cost is the production-relevant figure. Full data: concourse.agency.