28 Sep 2026 · 4 min read

Same spend, different contract: the real test of an AI media buyer

Two model-driven buyer agents committed the same amount of money to the same brief. One of them got a better deal. That finding, from Concourse Bench v1 (concourse.agency), matters more than any headline completion rate, because it is the question a media buyer or procurement auditor actually asks six weeks into a campaign: not whether the buy finished, but what the buy actually got you.

TL;DR. GPT-6 Astra and GPT-5.6 Sol both committed £18,706.89 against the same fictional UK retail brief. Their contract terms differed: make-good provisions and promised viewability floors diverged despite identical headline spend. Across 110 contracts produced in 55 completed buys, 95 promised viewability floors below 70 per cent. Completion is a threshold requirement. Contract quality is what determines whether a completed buy was worth completing.

What the simulation tested

Each buyer in Concourse Bench v1 operated as a model plus buyer harness system. The harness supplied the brief, checked every action, and recorded the full negotiation. Eight model-driven buyer agents ran a single fictional UK retail brief against five inventory packages and four prospective variants, with three fresh attempts per model per variant. Sellers A, B, and C held fixed software policies throughout: pricing floors and delivery terms did not bend to buyer pressure. Of 96 registered attempts, 91 were assessed; five technical interruptions were excluded. 55 of those 91 attempts ended in a completed contract. The full data and methodology are published at concourse.agency.

Completion is a threshold. What was contracted is the question.

GPT-6 Astra committed £18,706.89 from a budget of £19,528.69. Net forecast reach came to 140,856, calculated as 149,565 gross minus 8,709 for audience overlap, against a brief target of 137,156. Completed reach sat approximately 2.7 per cent above target; spend sat approximately 4.2 per cent below budget. By headline numbers, a clean buy.

GPT-5.6 Sol matched that figure exactly: £18,706.89 in committed spend. But the contract terms were not equivalent. Make-good provisions differed. Promised viewability floors differed. The four delivery guarantees across both buys were each set at 95 per cent, so the face of the deal looked identical. What the contracts actually said about what happened when delivery fell short, and what viewability standard was being promised, was not the same. A procurement team reviewing those two outputs at the invoice line would see the same number. A team reviewing the contract detail would face different risk exposures.

What scale looks like: 95 floors below 70 per cent across 110 contracts

The same-price, different-protection problem is not a one-pair anomaly. Across the 110 contracts produced in the 55 completed buys, 95 promised viewability floors below 70 per cent. A media plan that has completed is not a media plan that has delivered against viewability expectations. A buyer that completes buys efficiently but systematically accepts low viewability floors is producing cost-efficient waste. Teams evaluating model-driven buyers on completion and spend metrics alone will not see this.

The completion picture: context, not headline

Two models completed all 12 of their registered attempts: GPT-5.6 Sol and Claude Fable 5.1. Claude Opus 5 completed 10 of its 10 assessed attempts. The 60 per cent overall completion rate across all models reflects genuine spread: some models found alternative paths when seller constraints blocked one route to the required reach; others stalled. That spread is real and it matters for deployment. But completion rate is the threshold condition. Once a buyer can reliably complete, the question shifts to what it completes on.

Where buyers fail: negotiation, not comprehension

The 36 attempts that did not reach completion did not fail at comprehension. Those models had read the brief. Breakdown occurred when seller policy constraints prevented assembly of the required reach within budget, and the buyer could not find an alternative path. A subtler failure mode is partial completion: a recorded Claude Sonnet simulation produced net forecast reach of 124,511 against a brief requiring 137,156, a shortfall of approximately 9.2 per cent. The buy finished. It did not comply.

What teams evaluating AI buyers should actually measure

Concourse Bench v1, published at concourse.agency in September 2026, is the first structured evidence set to run this comparison across eight buyer models and multiple fresh attempts at the same brief. The evaluation criteria it points toward: completion rate (necessary, not sufficient); spend efficiency (necessary, not sufficient); and contract terms — make-good provisions, viewability floors, audience guarantees. Two models that produce the same spend and the same forecast reach are not interchangeable if one systematically accepts lower viewability floors or weaker make-good terms than the other. The finding that 95 of 110 viewability floors came in below 70 per cent is a finding about what completion alone measures. If your evaluation framework stops at completion, you are measuring the wrong thing.

Disclosure: All figures are sourced from Concourse Bench v1, published by Alkimi in September 2026. All seller interactions occurred in controlled recorded simulations against fixed fictional sellers with fixed software policies. No live publisher inventory or real media contracts are represented.

All articles

Speak to the team