28 Sep 2026 · 5 min read
How does an AI buyer fail a media brief?
Forty-one of 96 registered attempts in Concourse Bench v1 (concourse.agency) did not produce a completed buy. Understanding those failures is at least as important as understanding the 55 that succeeded. The benchmark distinguishes three failure modes, and the distinction matters for anyone trying to diagnose what needs to improve.
What are the three ways an AI buyer fails?
The first is budget overrun. The model negotiated individual contracts that were each valid but committed a total portfolio spend above the registered budget. This is a portfolio assembly failure: the model reached agreement with sellers but could not combine those agreements into a portfolio within the spend limit.
The second is missed brief requirement. The committed portfolio did not satisfy one or more requirements in the buying brief. Insufficient forecasted reach, wrong audience, wrong channel, or another specified condition not met. This is a buying quality failure: the model closed deals but not the right deals for the given brief.
The third is technical interruption. The run was stopped by a technical issue before it could be assessed. Five of the 96 registered attempts in v1 were unassessed for this reason. These five are reported separately, with the denominator kept visible, so that the interruption rate is not absorbed into a lower registered count and made invisible.
Why does the distinction between failure modes matter?
A budget overrun and a missed brief requirement point to different problems. A budget overrun suggests the model could negotiate individual contracts but struggled with the portfolio optimisation problem: combining agreements across three sellers into a valid total within the spend limit. A missed brief requirement suggests the model either misread the brief or deprioritised certain conditions when negotiating.
Both produce a failed run. They do not point to the same weakness, and they do not suggest the same remediation.
A technical interruption is a different category entirely. It is not a buying failure. It is an execution environment failure, and it should be reported separately because conflating it with the buying failure modes would overstate how often models fail for buying-related reasons and understate the role of infrastructure reliability.
What does a negotiation success without a portfolio completion look like?
The benchmark records a specific case where the model reached agreement with individual sellers on valid contracts but did not produce a portfolio satisfying the registered completion rule. The model negotiated successfully at the contract level and failed at the portfolio assembly level.
This is a distinct failure pattern worth naming precisely: successful individual negotiation, failed portfolio construction. A model can handle the counterparty interaction competently and still fail the final step, which requires optimising across the outcomes of three separate negotiations simultaneously, within a fixed budget, against the brief's requirements. These are different cognitive tasks, and they can fail independently.
Which models were affected, and in which situations?
The benchmark does not publish a per-model breakdown of failure mode in the v1 results. The five unassessed attempts included two Opus attempts and one Sonnet attempt, with the remainder distributed across the data. The distribution of budget overrun versus missed brief failures across the 36 assessed incomplete runs is not disaggregated in the publicly available tables.
What the situational data does show is that the firmer buying situation, where sellers held hard on pricing and terms, produced the most incomplete runs. Astra dropped from 3 of 3 in the reference situation to 1 of 3 in the firmer situation. Terra completed 2 of 3 in the reference situation and 0 of 3 in both the firmer and tighter situations. The most demanding conditions produced the most failures, which is an expected pattern, but the steepness of that drop varies significantly by model.
What does this mean for deployment evaluation?
A deployment test that measures only overall completion rate loses information that matters. Two models with a 9 of 12 completion rate might have completely different failure profiles: one failing because of budget overrun in the firmer situation, the other failing because of missed brief requirements in the tighter situation. The appropriate response to the first is different from the appropriate response to the second.
For practitioners designing an AI buying system or evaluating one, separating the failure modes at the start of the evaluation is more useful than collapsing them into a single pass/fail number. Budget overrun points to portfolio optimisation. Missed brief requirements point to brief parsing and negotiation strategy. Technical interruptions point to execution environment reliability. These are three separate improvement levers.
How does the firmer situation change the failure picture?
The firmer situation, where sellers held hard on terms, appears to have increased both budget overrun risk and the rate of incomplete portfolios. A seller that will not move on price puts pressure on the buyer's total budget across three negotiations. A model that manages the optimisation well under standard conditions may fail when there is less room to manoeuvre at the individual contract level.
This is the condition that most closely resembles actual trading pressure, where inventory is competitive and sellers hold the stronger position. Models evaluated only in the reference (standard baseline) situation would appear more reliable than they are when conditions tighten.
What the failure data does not establish
V1 is a controlled simulation, not a live-market study. Technical interruption rates in production may differ from benchmark conditions. Budget overrun and missed brief failures in a live context will depend on the quality of the brief, the range of inventory available, and the seller behaviour encountered. The three failure modes identified in v1 are useful as a taxonomy for evaluation design, not as production failure rate predictions.
All figures come from Concourse Bench v1, published by Alkimi at concourse.agency.