28 Sep 2026 · 4 min read
Three questions to ask before using AI media buying cost data
Cost comparisons for AI media-buying workflows are appearing in vendor briefings and trade coverage. Most of them are not wrong exactly. They are just not answering a useful question. The figure that tells you GPT-5.6 Sol costs $35.13 to complete a media buy is not the same figure as the one that tells you what Sol costs to run across your full buying workload. Before using any cost figure in a deployment decision, three questions determine whether it is usable.
TL;DR: Concourse Bench v1 (concourse.agency) tested eight AI buyer models across multiple attempts at the same brief. Sol's cost across all 12 workload attempts was $35.13. Fable's was $129.03. The full eight-model suite ran at $390.16. Those figures differ from per-completed-buy figures because they include failed attempts. Question 1: which population? Question 2: what happened to the failed runs? Question 3: what does the figure exclude?
Question 1: which population is this figure drawn from?
A cost per completed buy is a cost per successful event. A cost per workload attempt is a cost per job submitted, regardless of outcome. For a model that completes every attempt, these figures are the same. For a model that completes 60% of its attempts, they are not.
In Concourse Bench v1, GPT-5.6 Sol completed all 12 of its registered attempts. Its cost figure is identical whether you measure per completed buy or per workload attempt. Claude Luna completed 0 of its assessed attempts. A cost per completed buy for Luna is undefined; a cost per workload attempt is a cost for a system that produced no completed output. Two different figures; two different deployment questions. When a vendor presents a cost figure, the first question is what the denominator is. Completed buys, or all attempts?
Question 2: what happened to the failed runs, and were their costs counted?
A model that fails to complete a buy still consumed compute. In a production deployment, that cost does not disappear. It is real infrastructure spend against a buy that delivered nothing. A cost figure derived from completed buys only, that excludes the cost of all failed attempts, understates the true workload cost by the failure rate.
In Concourse Bench v1, the five technical interruptions across 96 attempts were excluded from assessment. That is a methodological choice: those runs failed for technical reasons outside model capability, and including their costs would have conflated infrastructure reliability with model performance. In a production deployment, you do not get that exclusion. Every failed run, technical interruption or otherwise, runs on your budget. The question is whether the cost benchmark you are reading made the same exclusions your production environment will not.
Question 3: what does the figure exclude beyond failed runs?
Compute cost is the most visible component and the most frequently cited. It is not the only cost in an AI buying workflow. Orchestration infrastructure, the buyer harness that checks actions and manages the negotiation loop, adds overhead that varies by harness design. Human review time, where any stage of the process requires a planner to check or approve, adds a cost that does not appear in API call logs. Integration costs, the overhead of connecting the AI buyer system to the relevant inventory systems and data sources, are typically excluded from benchmark figures entirely because they are environment-specific.
In Concourse Bench v1, the reported costs are compute costs for the model calls that ran the buying process. They do not include harness infrastructure costs, which were controlled and held constant. They do not include human review time, which was not part of the evaluated workflow. They do not include any integration costs. That is a clean benchmark for model comparison purposes. It is not a total cost of ownership figure for a production deployment.
The figures, read correctly
With those three questions answered, the Concourse Bench v1 cost figures are interpretable. Sol's $35.13 is its all-attempt cost across a 12-attempt workload in which it completed all 12. Fable's $129.03 is its all-attempt cost across a 12-attempt workload in which it completed all 12. Both figures are compute costs, harness costs excluded, human review costs excluded. The eight-model suite ran at $390.16 total across all registered attempts. These are model comparison figures. They are not production deployment cost estimates. Used as the former, they are useful. Used as the latter, they will produce a deployment cost that is lower than the reality. Full methodology at concourse.agency.