28 August 2026

How to Evaluate an Agentic Media Buying Platform: 7 Questions Every Buyer Should Ask

The right way to evaluate an agentic media buying platform is to ignore the autonomy pitch and interrogate the boundaries and the record. Every vendor will soon claim its agents negotiate well and act fast, and buyers will have no way to tell those claims apart. What separates a platform worth piloting from one worth avoiding is not how much the agent can do, but what it cannot do without permission, and whether it can prove what it did afterwards. These seven questions are built to surface that.

TL;DR. Ask, in order: where does human sign-off sit, is every decision auditable, is the deal recorded somewhere both sides trust, what are the guardrails, how does a pilot actually work, what does it integrate with, and how is success measured against a human baseline. A strong platform answers each with specifics. A weak one answers with adjectives. The questions are ordered by how much they reduce the buyer's personal risk, because in practice the person signing off on an agentic pilot is not optimising for the best campaign, they are trying not to be the one who put budget through something they could not later explain.

Why these questions, and why in this order?

Because the failure that ends careers in this category is not a mediocre campaign. It is an autonomous system that made a decision nobody sanctioned and nobody can account for, at a speed that turned a small error into a large one before anyone noticed. Analysts call it silent failure at scale: an agent doing exactly what it was told against a flawed signal, repeatedly, until the damage is done. Every question below is aimed at that failure, not at capability, because capability is what vendors compete on and accountability is what buyers get blamed for.

The order matters too. The first three questions, sign-off, auditability, and the shared record, are the ones that determine whether you can defend the decision in six months. The next two, guardrails and the pilot, determine whether you can contain the risk while you learn. The last two, integration and measurement, determine whether it will work in your stack and whether you can prove it did. Reduce the personal exposure first, then assess the performance.

Question 1: Where exactly does a human have to sign off?

This is the question that matters most, and the answer should be specific to the point of being a table. Which actions can the agent take entirely on its own, which require a human to approve first, and where is the line drawn? A credible platform draws that line by stakes: reversible, low-value actions run unsupervised, while anything that commits significant spend or opens new deal terms sits behind an approval gate. The current industry specifications build this in, requiring human approval on any path that commits spend above a value threshold.

Be suspicious of two answers. One is "the agent is fully autonomous," which means the interrupt that stops a compounding error has been removed. The other is a vague "there's always a human in the loop," which tells you nothing about where. The useful follow-up is: show me the threshold, and show me what happens above it. If the vendor cannot draw the line precisely, the line does not exist precisely, and neither does your control.

Question 2: Can any decision be traced to the reasoning behind it, months later?

An agent's value is its judgement across many levers at once. The cost of that judgement is that it is harder to reconstruct than a rule that either fired or did not. So the question is whether the platform records not just what the agent did but why, in a form a human can read long after the campaign closed.

The test is concrete. Ask the vendor to show you how you would find out, in month six, why the agent moved a third of the budget in week three. If the answer is a dashboard of outcomes and a natural-language summary with no recoverable logic between the decision and the result, the platform is not auditable in the sense you need. The standards bodies make the same point in stronger terms: a model narrating what it did in conversational prose is not a transaction record. You are not buying the ability to see what happened. You are buying the ability to explain it to someone who is questioning the spend, and a summary will not survive that conversation.

Question 3: Is the deal recorded somewhere the buyer and seller both trust?

This is the question most buyers do not think to ask, and it is the one that separates a durable platform from a fast one. When a buyer's agent and a seller's agent transact, each writes the deal into its own system by default. If those two records diverge, every later decision each agent makes rests on a different version of reality, and the divergence compounds silently until reconciliation, by which point the campaign is over.

The scale of this is not hypothetical. In a controlled simulation of 90,202 agentic transactions mapped to current industry specifications, two agents that had just agreed the same deal recorded its terms differently in 95.3% of cases, with both sides passing standard reconciliation checks throughout. That figure comes from a model of the specification architecture rather than a live system, and no platform has published its own equivalent, which is itself worth noting. Ask the vendor directly: do the buyer's agent and the seller's agent reference the same record, or does each keep its own, and when they disagree, which is authoritative and how is the disagreement surfaced before the campaign ends? A platform that gives both agents one shared record has solved a problem the others are quietly carrying, and it is a problem both competing standards bodies openly agree is not yet solved at scale.

Question 4: What are the guardrails, and are they deterministic?

Guardrails are the fixed limits inside which the agent operates: spend caps, approval thresholds, brand-safety requirements, and rules on what the agent may negotiate. The important property is that they are deterministic where the agent's reasoning is not. The agent's decisions are probabilistic and situational; the guardrails should be hard rules that do not bend under optimisation pressure.

This matters because an ungoverned agent optimises toward whatever it is measured on, including in ways a human never would. It can erode yield to maximise a fill rate, or misapply pricing tiers, at a speed that drains budget before review. Ask what the agent is prevented from doing regardless of what its optimisation suggests, and ask whether those limits are enforced deterministically or are themselves subject to the agent's judgement. Guardrails the agent can reason its way around are not guardrails.

Question 5: What does a pilot actually look like?

A serious platform expects a bounded, reversible pilot before it touches full budget control, and can describe one specifically. The pattern worth looking for is a proof period: the agent runs in recommend-only mode or under tight caps, long enough to demonstrate that its decisions hold up, and its autonomy widens only after reliability is shown. The operators worth taking seriously describe two tracks running in parallel, an experimental track under close oversight and a production track that scales only what has proven out. This graded approach, sometimes called earned autonomy, treats autonomy as something the agent earns against a record rather than a setting switched on at go-live.

The answer to avoid is one that pushes for broad autonomy fast, or that cannot articulate a proof period at all. Pilot-stage autonomy quietly becoming full budget control, with no demonstrated reliability in between, is the exact path to the failure the first four questions were built to prevent. Ask what the smallest safe starting point is, how long the proof period runs, and what specifically has to be true before the agent's control widens.

Question 6: What does it integrate with, and what does integration cost?

Agentic buying is only useful if the agent can reach the inventory, the data, and the reporting the buyer already relies on. Ask which standards the platform is built on, because the category is currently split between competing technical frameworks, and a platform aligned with the standard your partners adopt will interoperate where a proprietary one will not. Ask what it connects to on the reporting side, since an agent whose decisions cannot flow into the systems where spend is reconciled creates a new blind spot rather than removing one.

The honest version of this question includes cost and complexity. What does integration actually require of your team, what has to be built, and how long before the agent is operating against real inventory? A platform that cannot answer this concretely is either early enough that you are the integration test, or vague enough that the real cost will surface after you have committed.

Question 7: How is success measured, and against what baseline?

The metric has to map to what the buyer's own leadership cares about, not to what flatters the platform. An agent that improves a proxy metric while the outcome that matters stays flat has optimised the dashboard, not the business. Ask what the platform measures, and insist the measure connect to a real outcome: cost per genuine result, working media reaching real inventory, spend a human can account for.

The baseline is the part buyers most often skip. The right comparison is not the agent against nothing, but the agent against the human process it is replacing. Ask whether the platform can show the agent's performance against a human baseline on the same campaigns, because "the agent did well" means little without "compared to what." A platform confident in its agents will welcome that comparison. One that deflects it is asking you to take the improvement on faith.

Putting the seven together

Score a platform on how specifically it answers, not on how impressively. The strong ones respond to each question with a threshold, a mechanism, a proof period, a baseline. The weak ones respond with autonomy, intelligence, speed, and other words that describe the pitch rather than the controls. The gap between those two kinds of answer is the gap between a platform you can pilot with your name on it and one that will eventually leave you explaining a decision you did not make and cannot reconstruct.

The category is real and moving fast: live cross-platform agentic buys are running, holding companies are executing agent-to-agent deals, and the standards are maturing month by month. That said, the people writing those standards are candid that agentic buying is not yet running at scale, which is exactly why the evaluation has to be disciplined now, while the difference between a governed platform and an ungoverned one is easy to see and before a bad pilot makes it personal.


This article references the IAB Tech Lab's published agentic advertising specifications, industry analysis of agentic governance and pilot design, and published simulation research into agentic deal reconciliation conducted by Alkimi. The simulation models the current specification architecture and is not an assessment of any specific production platform.

Entering Alkimi Marketplace...