28 Sep 2026 · 6 min read
Why Alkimi built a benchmark before building a marketplace
When Alkimi began describing itself as an agentic marketplace for media trading, the natural question from any sophisticated buyer was: can AI actually buy media? Not in principle, not in a demo, not in a controlled environment where every variable is managed. In a real buying scenario, against real sellers, with a real brief and a real budget. We did not have a clean answer to that question. We suspected the answer was yes, in some conditions and for some models. But suspicion is not evidence, and building a marketplace on an untested claim is a liability when buyers do diligence.
Why did the claim need testing before it was made?
Marketing built on untested capability claims fails in two ways. The first is operational: when a buyer deploys a system that has been oversold, performance falls short, trust erodes, and attribution is murky. The second is reputational: when a claim is tested by others and found to be weak, the damage extends beyond the specific claim to the credibility of everything else the company says. In a market where AI capability claims are proliferating and where most of them are vague enough to be unfalsifiable, specificity is a competitive advantage. But specificity cuts both ways.
The alternative to testing your own claim before making it is to make the claim and hope that buyers do not probe hard enough to expose the gap. That is a short-term strategy with a predictable end. Alkimi's marketplace depends on buyers trusting that AI agents can transact on their behalf. That trust cannot be built on assertion. It has to be built on evidence.
What did the industry know about AI buying capability before the benchmark?
Before Concourse Bench v1 was published, there was no structured, multi-model, independently documented evidence of AI buying capability in a media context. Individual vendors had internal performance data, but that data was not published, not independently verified, and not comparable across systems. Buyers evaluating AI buying tools had to rely on demos, vendor claims, and case studies that were selected for success. None of that is a basis for confident procurement decisions.
The practical consequence was that buyers could not answer basic questions about AI buying readiness: which models can complete buying tasks reliably? What goes wrong when they fail? Are the deals they commit to contractually sound? These are not edge-case questions. They are the questions any responsible buyer would ask before deploying an AI agent on real budget.
What does Concourse Bench v1 actually measure?
The benchmark runs AI models through structured media buying tasks. Each task presents the model with a defined brief: a budget, an audience specification, channel requirements, quality parameters, and timing constraints. The model must negotiate with sellers and assemble a portfolio that satisfies the brief. Completion requires producing a deal record that is budget-compliant, brief-satisfying, and contractually coherent.
The benchmark produces three categories of output. The first is a completion rate: out of the total attempts made, what proportion resulted in a completed buy. The second is a failure mode breakdown: for the attempts that did not complete, what failed and at which stage. The third is contractual quality data: for the deals that were committed to, how close were they to the brief's requirements and how sound were the contract terms.
Of 96 registered attempts, 55 produced a completed buy. 41 did not. The 41 failures broke down across three distinct failure modes, each pointing to different weaknesses in different models. That breakdown is more useful than an aggregate pass rate because it tells you what to fix, not just that something is broken.
Why publish a benchmark that tests models other than your own?
The decision to benchmark multiple models, rather than running internal tests and publishing summary conclusions, was deliberate. An internal test of Alkimi's own infrastructure would answer the question of whether our system works. It would not answer the question of whether AI buying works. Those are different questions, and the second one is the one buyers need answered before they will trust the first.
A multi-model benchmark is also more credible. If we had published data showing only that models which work well on our platform perform well, the finding would be circular. By running multiple models through the same structured tasks, the benchmark produces comparative data that buyers can use to evaluate model selection independently of platform choice. The benchmark is useful to the industry whether or not buyers choose Alkimi as their marketplace.
There is also a commercial rationale. Alkimi's marketplace depends on AI buyers being real. If AI buying does not work, the marketplace has no demand side. It is in Alkimi's direct interest to demonstrate that AI buying works, what conditions it requires, and which models are currently best placed to do it. The benchmark serves that interest by producing evidence rather than assertion.
What does the benchmark enable in terms of content and positioning?
The benchmark is the evidence base for every substantive claim Alkimi makes about AI buying. When we say that AI agents can complete media buying tasks, we can point to 55 completions across 96 attempts. When we say that completion rate is the minimum threshold for autonomous deployment, we can point to the specific models that achieved it and the specific ones that did not. When we say that failure mode analysis is more useful than aggregate pass rates, we can show the three distinct failure modes and what each one implies for remediation.
Without the benchmark, these would be opinions. With it, they are findings. The difference between an opinion and a finding is the difference between a vendor claim and a procurement-relevant piece of evidence. Buyers doing diligence on AI buying tools can engage with findings in a way they cannot engage with opinions.
What does the benchmark reveal about the current state of AI buying?
The headline finding is straightforward: AI buying works, and it does not always work, and the failure modes are specific and diagnosable. A 57% completion rate across 96 attempts is not a pass. It is not a fail either. It is evidence of a technology that has real capability and real limitations, at a specific moment in time, under specific conditions.
That is exactly what the industry needed before the market for AI buying tools matured past the point where honest assessment was possible. Alkimi built the benchmark before building the marketplace because the marketplace depends on buyers having realistic expectations, grounded in evidence, about what AI buying can and cannot do. The evidence is now available. The marketplace is built on it.