28 Sep 2026 · 7 min read
Three ways an AI buyer fails a media brief, and what each failure tells you
Of the 96 attempts registered in Concourse Bench v1, 55 produced a completed buy. 41 did not. A headline failure rate of 43% would be sufficient to raise serious questions about the readiness of AI buyers for autonomous deployment. But the headline number is the least useful part of the data. The 41 failures did not all fail in the same way. They fell into three distinct failure modes, and the distinction between those failure modes determines what needs to change, in which models, at which stage of the buying process. A single pass/fail rate cannot tell you that.
What is a budget overrun failure, and what does it indicate?
A budget overrun failure occurs when an AI buyer negotiates individual contracts that are each valid in isolation, but whose combined total exceeds the brief's stated budget. The model closed deals. It produced a portfolio. But the portfolio's total spend commitment is above the budget it was given to work with.
This is a portfolio assembly failure, not a negotiation failure. The model demonstrated that it could conduct individual negotiations and reach agreement with sellers. What it failed to do was maintain an accurate running total of its portfolio-level spend commitment across all negotiations simultaneously. At some point during the process, the model committed to a deal that tipped the portfolio over budget, either because it lost track of the cumulative total or because it processed the individual deal as if the full budget were still available.
The implication for diagnosis is specific: this model has individual negotiation capability but not portfolio budget tracking capability. The improvement target is not negotiation strategy. It is the mechanism by which the model maintains and applies a cumulative budget constraint across parallel negotiations. Training approaches that address this specifically (reinforcing budget constraint propagation across concurrent negotiation threads) are more likely to fix the problem than general negotiation quality improvements.
Budget overrun failures are also the most commercially immediate. A campaign that ran over budget by 20% because an AI buyer lost track of its cumulative commitments creates a real cost to the buyer and a real contractual obligation to the sellers. The failure is not just a quality gap. It is a financial liability.
What is a missed brief requirement failure, and what does it indicate?
A missed brief requirement failure occurs when the committed portfolio does not satisfy one or more requirements stated in the buying brief. The brief specified an audience, a channel, a minimum viewability floor, or some other quality parameter. The portfolio that was committed does not meet that specification.
This failure mode has two distinct origins, and distinguishing between them requires inspecting the negotiation record rather than just the final portfolio. The first origin is brief adherence failure under seller pressure: the model accepted deal terms outside the brief's requirements during negotiation, either because a seller proposed alternative terms and the model accepted them without checking against the brief, or because the model deprioritised brief compliance in order to close a deal. The second origin is coverage failure: the model negotiated correctly but failed to reach agreement with any seller who could satisfy a specific brief requirement, leaving a gap in the portfolio that cannot be filled.
These two origins require different remediation. Brief adherence failure under pressure points to a model that needs stronger constraint maintenance during negotiation: it should be more resistant to seller proposals that would violate the brief, and more willing to walk away from a negotiation that cannot be satisfied within the brief's parameters. Coverage failure points to a model that needs better seller selection and engagement strategy: it pursued the wrong sellers, or pursued the right sellers with positions that could not reach agreement.
For brand safety purposes, missed brief requirement failures are particularly significant. A brief that specifies content adjacency requirements or audience exclusions is setting parameters that protect the advertiser's brand. A model that commits to deals outside those parameters, even if the individual deals are otherwise valid, has failed to maintain the brand safety conditions under which it was authorised to operate. The financial consequences may be limited. The reputational consequences may not be.
What is a technical interruption, and why is it reported separately?
Five of the 96 attempts in Concourse Bench v1 were stopped by a technical issue before assessment was possible. The model did not complete the buying task, but the reason was not a buying failure. The task was interrupted before it reached a conclusion.
Technical interruptions are reported as a separate category rather than included in the failure rate because conflating them with buying failures would misrepresent both. An AI buyer that failed because its negotiation strategy was poor is a different problem from an AI buyer that was interrupted by an environment failure before it could demonstrate its negotiation capability. The first points to model improvement. The second points to infrastructure reliability improvement. Aggregating them into a single failure count obscures which lever needs pulling.
The practical significance of separating technical interruptions from buying failures is that it preserves the interpretability of the buying failure data. If technical interruptions are included in the failure count, a model with a low completion rate might appear weaker than it is because of infrastructure issues rather than buying capability gaps. Conversely, a model with good buying capability but a fragile execution environment would be penalised twice: once for the technical interruptions and once for the reduced sample size reducing confidence in its buying performance data.
Why does the distinction between failure modes matter for deployment decisions?
A buyer considering deploying an AI buyer needs to understand not just whether a model passes or fails a benchmark, but which failure mode it is most likely to exhibit when it does fail. That information is not available from a completion rate alone.
For buyers running campaigns with tight budgets and strict financial controls, a model with a history of budget overrun failures is a higher risk than one with brief compliance failures. The budget overrun failure is a financial liability that creates contractual obligations. The brief compliance failure, depending on the specific requirement missed, may be more manageable.
For buyers running brand-sensitive campaigns with specific content adjacency requirements, a model with brief compliance failures (particularly of the brief adherence under pressure variant) is a higher risk than one with budget overrun failures. The brand safety implication of a compliance failure may be more consequential than the financial implication of a budget overrun.
Neither failure mode is acceptable in production at scale. Both need to be remediated before a model is ready for autonomous deployment against real client budgets. But knowing which failure mode dominates in a specific model tells you what to work on, and knowing which failure mode is most relevant to your specific campaign type tells you which model risks are most material for your context.
What should buyers take from the Concourse data before making deployment decisions?
Three things. First, the completion rate establishes the floor: a model with a very low completion rate is not ready for autonomous deployment, regardless of the quality of the completions it does produce. Second, the failure mode breakdown tells you where the model's capability gaps are and whether those gaps are material to your specific buying context. Third, the technical interruption count tells you something about the reliability of the execution environment, which is a separate question from the model's buying capability.
The most important implication of the failure mode analysis is one that applies regardless of which specific model you are evaluating: before deploying an AI buyer against a real brief with a real budget, you need to know which type of failure is most likely in your environment. Because the fix is not the same. A model that loses track of portfolio budgets needs different improvement from a model that drifts from brief requirements under counterparty pressure. Understanding which problem you have is the starting point for solving it. The Concourse data provides exactly that understanding, in a form that is transferable across model and deployment decisions.