28 Sep 2026 · 4 min read
Which AI models can reliably close a media deal?
Not all AI models perform equally when given a media-buying task. Concourse Bench v1 (concourse.agency) provides the first structured, multi-model evidence on this question. Eight model-driven buyers each ran four buying situations with three registered attempts per situation, for 96 registered attempts in total. Five attempts were unassessed because of technical interruptions; the remaining 91 were fully evaluated.
The results divide the eight models into three groups: those that completed reliably, those that completed inconsistently, and those that completed nothing.
Which models completed reliably?
Sol and Fable each completed all 12 of their 12 registered attempts, across all four buying situations: reference (standard baseline), tighter (compressed budget), choice (wider inventory pool), and firmer (sellers holding hard on terms). No other model in v1 matched this record.
Opus completed 10 of its 12 registered attempts. Two Opus attempts were unassessed because of technical interruptions; among the 10 assessed attempts, all 10 completed. Whether the two unassessed attempts would have completed is not established by the data.
Astra completed 9 of 12 registered attempts. Sonnet completed 9 of 12, with one additional interruption in the reference situation.
Which models completed inconsistently?
Terra completed 3 of 12 registered attempts: 2 of 3 in the reference situation, 0 of 3 in the tighter situation, 1 of 3 in the choice situation, and 0 of 3 in the firmer situation. Terra's completions were confined to the less demanding buying conditions. When the budget was compressed or sellers held firm, Terra produced no completed buys.
The pattern in Terra's results is worth examining closely. A model that completes the standard task but fails when conditions tighten is not a model that can be relied upon in real buying environments, where standard conditions are the exception rather than the rule. Buyers do not usually face only the easiest version of the task.
Which models completed nothing?
Luna and Haiku each completed 0 of 12 registered attempts, across all four buying situations. Luna spent $0.15 in known API costs across all 12 attempts. Haiku spent $0.13. Both models ran every attempt, incurred cost, and produced no completed buys.
The benchmark notes that cost per success is undefined for a model with no completions. This is not just an accounting point. A zero-completion model cannot be evaluated on cost efficiency because there is no output against which to measure the cost. A model with zero completions cannot be deployed as an autonomous media buyer regardless of its performance on other tasks.
How do the buying situations affect completion rates?
The firmer situation, where sellers held hard on pricing and terms, was the most difficult across the board. Astra dropped from 3 of 3 in the reference situation to 1 of 3 when sellers held firm. Terra, which managed 2 of 3 in the reference situation, completed 0 of 3 in the firmer situation. Even the highest-performing models, Sol and Fable, had to maintain consistent performance across the firmer situation to achieve 12 of 12.
The tighter budget situation was the second most demanding. Terra completed 0 of 3 attempts when the budget was compressed, and several other models saw their completion rates decline under tighter conditions. These are the conditions that most closely resemble constrained real-world buying: a brief to meet, a budget that does not stretch, and sellers who will not easily concede.
What completion rate is needed for deployment?
The benchmark does not set a threshold, and v1 does not include a human baseline. The question of what completion rate is acceptable for a given deployment depends on the specific use case: the volume of buys, the tolerance for manual review, and the cost of an incomplete run relative to the cost of human oversight.
What v1 establishes is the floor: a model with zero completions offers no viable deployment path as a buyer. Beyond that, the threshold is a practitioner decision. A 9 of 12 rate differs meaningfully from 12 of 12, and whether that difference matters depends on the deployment context and what happens to incomplete runs in practice.
What do the sellers reveal about these results?
The v1 sellers were fixed software policies, not language models. Model-driven sellers are planned for a later version of the benchmark. The completion rates above therefore describe AI buyer performance against deterministic, rule-based counterparts. How completion rates shift when the seller is also a model is a question the benchmark has not yet answered.
What the completion table cannot tell you
Completion is a necessary condition for deployment, not a sufficient one. A model that completes every attempt while accepting weak contractual terms, committing above-target spend on individual contracts, or missing brief requirements is completing the workflow without doing the job. V1 records both completion and contractual quality, and the viewability data shows a significant gap between the two: 95 of the 110 contracts in completed buys promised a viewability floor below 70%.
The completion table is the starting point for an evaluation, not the end point. A model that sits at the top of the completion table and the bottom of the contractual quality data is not a good buyer. Both dimensions need measuring.
All figures come from Concourse Bench v1, published by Alkimi at concourse.agency. The benchmark represents controlled simulation, not live market performance.