28 Sep 2026 · 4 min read
Can an AI agent actually complete a media buy?
Fifty-five of 96 registered attempts ended with a completed buy. That headline has circulated since Concourse Bench v1 was published in September 2026. On its own, it says almost nothing useful. It becomes useful only when you look at which models produced those 55 completions, and which produced none.
Concourse Bench v1 (concourse.agency), published by Alkimi, placed eight model-driven buyers into a controlled simulated market. Each received a fictional media-buying brief, one budget, and three sellers to negotiate with. The sellers were fixed software policies, not language models, each running deterministic behaviour. The benchmark recorded whether the system completed the portfolio, what it committed, and how the model performed across four buying conditions.
What did the benchmark test?
Four buying situations tested the same underlying brief in different forms. The reference situation was a standard baseline. The tighter situation compressed the available budget. The choice situation widened the inventory pool. The firmer situation had sellers holding harder on pricing and terms. Each model-driven buyer ran three registered attempts per situation, giving 96 registered attempts in total. Five attempts were unassessed because of technical interruptions; the remaining 91 were fully evaluated.
Completion required the buyer to negotiate with the three sellers and commit a portfolio within the registered rules: total spend within budget, brief requirements met. The benchmark evaluated the model as part of a system. The buyer harness, the software supplying the brief and checking every action, was constant across all runs. Completion figures describe the system, not the model in isolation.
Which models completed the buy, and how often?
Sol and Fable each completed all 12 of their 12 registered attempts, across all four buying situations. Opus completed 10 of its 12 registered attempts (10 of the 10 that were assessed; two were unassessed due to technical interruptions). Astra completed 9 of 12. Sonnet completed 9 of 12. Terra completed 3 of 12. Luna and Haiku each completed 0 of 12.
The spread from 12 of 12 to 0 of 12 is not a marginal performance difference. Luna and Haiku completed nothing across all four buying situations and all three attempts per situation. A model with zero completions cannot serve as a media buyer regardless of its performance on other tasks. Sol and Fable cleared the completion bar every time. Everything between those poles requires more specific scrutiny.
Where did completions break down?
The firmer situation, where sellers held hard on their terms, produced the most failures. Astra dropped from 3 of 3 in the reference situation to 1 of 3 when sellers held firm. Terra, which completed 2 of 3 in the reference situation, completed 0 of 3 in both the firmer and tighter situations. The tighter budget situation was the second most difficult across the board.
This pattern matters beyond the table. Real buying conditions are not uniformly standard. A model that performs well when sellers are accommodating and budgets are comfortable may fail in precisely the situations where completion matters most, when margins are tight or the seller's position does not move.
What does completion not establish?
The benchmark is explicit on this. Completion is the first layer, not the whole picture. A portfolio can be completed while containing weak contractual protections, above-target spend on individual contracts, or terms that fall short of the brief's requirements.
The viewability data from v1 makes this concrete. Across 110 contracts in the 55 completed buys, 95 promised a viewability floor below 70%. Fifteen promised 70% or above. An agent that completes every run while accepting poor contractual terms has finished the workflow and missed the brief. Concourse measures completion and contractual quality separately, because those are two different scores.
What about the sellers?
In v1, the three sellers were fixed software policies, not AI models. Model-driven sellers are planned for a later version of the benchmark. This means v1 measures AI buyer performance against deterministic counterparts, not against other AI agents. The results should be read in that context. How completion rates shift when the seller is also a model is a different question, with a different answer.
What the data supports and what it does not
The figures throughout this piece come from Concourse Bench v1, published by Alkimi at concourse.agency. The benchmark represents controlled simulation, not live market performance. API costs are recorded model API expense, not full operating costs. There is no human baseline yet; one is planned, and as the benchmark states, it is the next thing needed before completion rates can be read as good or bad rather than simply recorded.
V1 is the first structured evidence on a question that has previously been answered mostly by demo and declaration. It establishes which models, in controlled conditions, can finish the job. It establishes which cannot. And it establishes that the answer is not the same for every model, every situation, or every kind of task within the workflow.
What is the useful question for practitioners?
The question that matters for anyone considering a deployment is not whether AI can buy media. Some of the models tested can. The question is which model-harness combination, under which buying conditions, meets the completion threshold required for the specific deployment in view. V1 provides a first layer of evidence. The next layers require a human baseline, live-market conditions, and data on AI-to-AI negotiation that later versions of the benchmark will begin to supply.