28 Sep 2026 · 4 min read
Why completing a media buy is not the same as doing a good one
The 55 completions from Concourse Bench v1 (concourse.agency) look reasonable until you examine what those completions actually committed to. Fifty-five of 96 registered attempts produced a finished portfolio. Inside those portfolios were 110 individual contracts. Ninety-five of them promised a viewability floor below 70%. Fifteen promised 70% or above.
That figure is the sharpest finding in v1, and it is the finding most likely to be missed if the benchmark is read only as a completion study.
What is a viewability floor, and why does it matter in a contract?
A viewability floor is the minimum viewability level the buyer secured in writing. It is a contractual promise from the seller, distinct from the viewability the inventory was advertised at and distinct from what is actually delivered during a campaign. In the context of Concourse Bench v1, published by Alkimi, it is the level the AI buyer negotiated and committed to in the deal record.
Viewability floors below 70% are not automatically meaningless, but 70% is a commonly referenced minimum in programmatic trading standards. A floor below that threshold means the buyer has accepted inventory with comparatively weak contractual protection. Across the 55 completed buys in v1, 86% of the contracts fell below that threshold. The agents completed the workflow and accepted weaker terms than the brief required.
Why did the agents accept weak viewability terms?
The benchmark records actions, not the model's private reasoning, so a definitive answer is not available from the data. What the data shows is that completion, as a success criterion, does not require strong viewability commitments. An agent optimising to complete the portfolio within budget can do so by accepting terms that a skilled human buyer would refuse or push harder to improve.
This is not unique to AI buying. Human buyers operating under time pressure, or with inadequate brief enforcement, can produce the same outcome. What the benchmark makes visible is the gap between "the portfolio is complete" and "the portfolio is good." In the v1 data, that gap is wide. A pass/fail completion score would hide it entirely.
Does completion rate correlate with viewability protection?
The benchmark does not publish a per-model breakdown of viewability terms in the v1 results. The 110 contracts figure is aggregate across all completed buys in the study, not separated by model. It is not possible from v1 alone to determine whether Sol and Fable, both of which completed 12 of 12 attempts, secured stronger viewability floors than Astra (9 of 12) or Sonnet (9 of 12). The viewability finding is a systemic pattern across the whole dataset, not a finding about any individual model.
That distinction matters for how the data should be used. A high completion rate does not certify strong contractual quality. The two measures are independent, and v1 establishes both.
What does a good buy look like, as opposed to a completed one?
Concourse Bench v1 illustrates the gap with a direct comparison: two portfolios committing identical cash, with the same forecasted delivery, but different make-good terms and different viewability floors. One is the better deal. Completion rates alone do not identify which.
The benchmark uses this comparison to make the case for its own next phase. V2 is described as "a stricter test of buying quality against fixed sellers." The implication is that completion was the right first question but not the last. An AI buyer that cannot complete the portfolio is not deployable. An AI buyer that completes the portfolio while systematically accepting poor terms is also not deployable, for different reasons.
How should practitioners read the viewability data?
The figures come from Concourse Bench v1, a controlled simulation published by Alkimi at concourse.agency. The sellers were fixed software policies, not live inventory sources. The viewability floors are negotiated contractual promises inside the simulation, not campaign delivery outcomes. The benchmark is not a production study, and the figures should not be read as predictions of what a deployed AI buyer would commit to in live trading.
What the data establishes is a proof of concept for a specific problem: completion metrics, measured alone, can conceal systematic failures in buying quality. If a benchmark, a deployment test, or a vendor evaluation measures only whether the AI finished the workflow, it is not measuring whether the AI did the job well. The v1 data puts a number on the size of that gap: 95 contracts out of 110.
What comes next in the benchmark?
Concourse v2 is described as a stricter test of buying quality, using different tasks and rules, making it a new benchmark rather than a rerun of v1. The separation between completion and quality that v1 makes visible becomes the explicit subject of v2. There is also a planned human baseline, which would allow quality measures to be compared against what a human buyer secures under the same conditions. Until that baseline exists, the viewability finding tells practitioners what the gap looks like, but not how deep it runs relative to human performance.
The point the v1 data earns is simple: a pass/fail completion number is necessary evidence, and not all models can produce it. But it is the first question, not the last. The viewability finding suggests the second question is the harder one.