28 Sep 2026 · 4 min read

How to evaluate an AI media buyer: what Concourse v1 reveals about the right methodology

Before Concourse Bench v1 (concourse.agency), there was no agreed methodology for evaluating AI media buyers. Vendors published internal figures. Agencies ran informal tests. Results were not comparable. Concourse v1 was an attempt to build something replicable. The methodology is at least as important as the findings it produced, because it is the methodology that will be challenged, extended, and improved in subsequent benchmarks. Here is what it established and why it matters.

TL;DR: Concourse Bench v1 tested eight model-driven buyer systems across 96 registered attempts against a single fictional UK retail brief with fixed sellers. The five design decisions that produced credible results were: fix the sellers, use multiple buying situations, run multiple fresh attempts per model per situation, measure completion and contract quality separately, and report failures by type. Any evaluation methodology that omits these will produce figures that cannot be compared across vendors or deployments.

Design decision 1: fix the sellers

In Concourse v1, Sellers A, B, and C operated on fixed software policies throughout the benchmark period. Pricing floors and delivery terms did not adjust to buyer pressure. This is the most consequential design decision in the benchmark. Without fixed sellers, a buyer that achieves better terms may be doing so because it found a more accommodating seller, not because it is a better negotiator. Fixed seller policies mean that every difference in outcome between buyer systems reflects genuine capability differences in the buying agent. If you are designing an evaluation and the sellers can respond differently to different buyers, the results are not comparable.

Design decision 2: use multiple buying situations

Concourse v1 tested four variants of the same brief: a reference condition and three progressively tighter or more constrained conditions (tighter-seller, choice-constrained, firmer-seller). A buyer that completes a reference buy reliably may fail under firmer-seller conditions; a buyer that fails at reference may not be worth testing further. Multiple situations reveal whether performance is robust or situation-specific. A single-condition evaluation tells you how a buyer performs in one context. Multiple conditions tell you whether it has a range.

In the firmer-seller condition, one buyer produced valid contracts and returned a completion state but missed the net reach target by 9.2%. That failure mode did not appear in the reference condition. It appeared when the buying environment became more constrained. Single-condition evaluations would not have found it.

Design decision 3: run multiple fresh attempts per model per situation

Each model ran three attempts per situation variant. Fresh attempts, not replays of the same session. This matters because AI buyer behaviour is not deterministic. A model that completes a buy on its first attempt and fails on its second and third is a different deployment proposition from a model that completes all three. Concourse v1 registered 96 attempts across eight models and four situations: three attempts per model per situation. That structure allowed reliability to be measured, not only peak performance.

Design decision 4: measure completion and contract quality separately

Completion rate records whether a buy finished. Contract quality records what the buy actually committed. Concourse v1 tracked both. The result was that some models with high completion rates produced contracts with systematically low viewability floors. Across all 55 completed buys, 95 of 110 contracts carried viewability floors below 70%. That is a contract quality finding, not a completion finding. An evaluation framework that measures only completion cannot produce it. Any buyer assessment that does not evaluate contract terms alongside completion rate is measuring whether the buy finished, not whether the buy was good.

Design decision 5: report failures by type

Concourse v1 distinguished between three failure types: technical interruptions (five cases, excluded from assessment), negotiation failures where seller constraints blocked assembly of the required reach, and partial completions where contracts were signed but the brief was not satisfied. These are different problems with different remedies. A technical interruption is an infrastructure problem. A negotiation failure is a model capability problem in constrained environments. A partial completion is an evaluation problem: the system reported done without checking brief compliance. Aggregating these into a single failure count obscures what needs to be fixed.

The four dimensions any evaluation should cover

Concourse v1 points toward four dimensions that any AI buyer evaluation should report separately: execution (did the buy complete against the brief requirements, not just the deal requirements); commercial quality (what were the contract terms, including viewability floors, make-good provisions, and audience guarantees); reliability (what was the completion rate across multiple fresh attempts at the same conditions, not only peak performance); and measured efficiency (what was the all-attempt cost, including failed runs, not only the per-completed-buy cost). The full methodology and data are published at concourse.agency.

All articles

Speak to the team