31 August 2026

Natural-Language Confirmation Is Not a Transaction: The Verification Problem at the Heart of Agentic Advertising

IAB Tech Lab's Shailley Singh: agentic advertising cannot rely on natural-language confirmation. Research across 90,202 simulated deals found 95.3% disagreement between separate agent records, falling to 0.19% with a shared record.

When an AI agent tells you it bought your media, that is not a receipt. That is a claim. The people writing the standards for agentic advertising have said so explicitly, and the structural reason is that natural-language output from a language model cannot stand as a deterministic record of a transaction.

TL;DR. Shailley Singh of IAB Tech Lab described the failure mode plainly in August 2026: agentic advertising cannot rely on "natural-language confirmation" from a model. Brian O'Kelley of the Ad Context Protocol said a deterministic, auditable record is a precondition, not a nice-to-have. Research modelled across 90,202 simulated transactions found a 95.3% disagreement rate between agents keeping separate records, falling to 0.19% with a shared record. The gap between "an agent said it happened" and "both sides hold the same account of what happened" is where a large fraction of agentic advertising investment will be lost if the verification layer is not built first.


What does "natural-language confirmation" actually mean?

The phrase came from Shailley Singh, head of standards at IAB Tech Lab, writing on the record to ADOTAT in August 2026. The architecture of agentic advertising, he said, cannot rely on "natural-language confirmation" from a model. A transaction record has to be deterministic, auditable, and one that both parties can reference.

Natural-language confirmation sounds like a technical term. It is not. It is the polite name for an agent telling you in English that it bought something.

Every demo in the agentic advertising space this year has included one. An agent reads a brief, builds a plan, and narrates the result in fluid, confident prose: it bought 4.2 million impressions across three SSPs, accepted the 30-day guarantee terms, locked the CPM at a specified rate. Singh's position, shared independently by Brian O'Kelley of the Ad Context Protocol in the same investigation, is that this narration is not the transaction. It is a description of one. The difference is structural, not cosmetic.

Why can a language model not confirm its own transaction?

The failure mode is not that models fabricate. It is that natural-language output is inherently ambiguous, and two systems parsing the same words will not necessarily arrive at the same structured meaning.

Singh named this explicitly. His concern was "differences between LLMs degrading interpretation of the transactional context." A buyer's agent and a seller's agent might agree, in the conversational sense, on the terms of a deal. When each walks away and writes down what was agreed, the two records will not necessarily match, because the same words can be parsed into different representations by different models.

Research published with WPP tested this directly. Modelled across 90,202 simulated agent-to-agent transactions, it found that under separate record-keeping, where each agent maintained its own account of the deal, the two records disagreed on at least one term in 95.3% of cases. The same transactions run with a single shared record that both agents wrote to produced a disagreement rate of 0.19%. The full methodology and codebase are available from WPP Research.

What does a deterministic record actually require?

Singh was precise about what the alternative looks like. The architecture needs "structured transaction objects, explicit transaction states, and systems of record that minimise hallucination or misinterpretation of context and that both sides can reference and audit." He added that existing standards like OpenDirect provide a foundation for "defining canonical transaction states rather than relying on natural-language confirmation."

O'Kelley, answering independently, used a single word when asked whether a shared machine-readable semantic layer is a precondition for cross-company agent buying: "Yes."

The point is not that the infrastructure needs to be built from scratch. It is that the natural-language layer currently used in demos is not the same as the structured-object layer that transaction verification requires. A system can produce fluent narration and still have no deterministic output behind it.

Where does the money go if the verification layer is missing?

The problem is invisible from inside either system, which is what makes it costly. In the WPP Research simulation, run to ninety days at holding-company scale, 677 million impressions settled with no agreed record of what had been delivered. Buyers paid against their numbers. Sellers invoiced against theirs. Neither flagged an error, because from inside each system, the account was internally consistent.

The simulation modelled holding-company scale, not the current live volume of the market, which Andrew Mole of pubX put at $2 to $3 thousand dollars a day in gross spend in an August 2026 interview with ADOTAT. But the mechanism is the same regardless of scale: when neither side has a shared record, each side's own account of the deal is, by default, correct. Disputes have no fact to resolve against.

What should buyers look for in a vendor's architecture?

The checklist from both standards bodies is the same one a buyer should apply. Does the platform produce structured transaction objects, not narrative summaries? Are transaction states explicit and machine-readable? Is there a system of record that both sides can query independently and receive the same answer?

An agent that narrates its buys confidently is not, by itself, evidence of any of these. The question to ask is not "did the agent buy the media." It is "where is that written down, and can you show me both sides of the ledger."

Singh put it plainly when describing what the architecture must provide: systems of record that "both sides can reference and audit." O'Kelley used the same phrase when asked about preconditions: "Yes." The two most competitive organisations in agentic advertising standards agree on what is required and agree it is not yet running at scale. The current phase of the market is small enough that this is still a design-phase problem. The decisions being made now will determine whether the category builds on a foundation that can hold.

Entering Alkimi Marketplace...