8 September 2026 · Updated 9 September 2026

Can AI Agents Be Trusted with Advertising Budgets? The Case For Defined Thresholds

AI agents are now making buying decisions across programmatic advertising at meaningful scale. The question "can agents be trusted with advertising budgets?" has a specific answer, and it is not yes or no.


By Alkimi

AI agents are now making buying decisions across programmatic advertising at meaningful scale. The question "can agents be trusted with advertising budgets?" has a specific answer, and it is not yes or no. It is: trusted to do what, up to what value, under what conditions? Trust in agentic advertising is not a binary state. It is a function of the approval architecture, and buyers who treat it otherwise will either constrain their agents too tightly to be useful or expose themselves to the category of failure that cannot be recovered from cleanly.

TL;DR: The trust question for agentic advertising is not about whether an AI agent is intelligent enough to manage a media budget. It is about whether the system around the agent has the right controls: defined spend thresholds, a shared deal record, pre-approval for material decisions, and a way to expand or contract agent authority as performance evidence builds. Buyers who have those controls in place can extend meaningful autonomy to agents. Buyers who do not are not managing risk, they are deferring it.

Why the trust question is being asked in the wrong way Most public discussion of AI trust in advertising frames the question as a capability problem. Can the agent make good decisions? Is the algorithm smart enough? Does the model understand brand nuance?

These are real questions, but they are the wrong starting point for a buyer assessing operational readiness. The capability of the underlying model is largely determined by the platform vendor and not directly in the buyer's control. What is in the buyer's control is the structure of the mandate, the thresholds that determine what the agent can do without human approval, and the audit infrastructure that shows what the agent actually did.

Trust, in this framing, is not a property of the agent. It is a property of the system the agent operates within. An agent with weak controls cannot be trusted regardless of its capability. An agent with strong controls can be extended significant autonomy even if its underlying model is imperfect, because the controls will catch and contain failures before they become material.

What a threshold architecture actually looks like A threshold architecture for a media buying agent defines, explicitly, which decisions the agent can make autonomously, which decisions require a human to approve before the agent acts, and which decisions are outside the agent's scope entirely.

The simplest version has three tiers. In the first tier, the agent acts autonomously within pre-approved parameters: buying against a defined audience on approved inventory categories, up to a daily spend cap. No human approval is needed for individual buys. In the second tier, the agent proposes a decision that exceeds a defined threshold (a deal above a certain value, a new inventory category not previously approved, a targeting signal outside the standard set) and waits for human approval before acting. In the third tier, certain decisions are reserved for humans regardless of whether the agent could technically execute them: mandate changes, new publisher relationships, decisions that involve novel data types.

IAB Tech Lab's Advertising Agent Messaging Protocol (AAMP) incorporates exactly this kind of tiered approval structure into its technical specification. The protocol defines decision types and the approval requirements for each, so that compliant buyer and seller systems can exchange approval states as part of the deal record rather than handling them through separate workflows that are easy to skip.

The earned autonomy model The tiered threshold architecture above describes a static state of trust: the agent operates with a defined scope and cannot exceed it without approval. But most deployments want a path to expanding agent authority over time, as the agent demonstrates reliable performance.

The earned autonomy model provides that path. It is a five-stage progression: observe (the agent watches decisions but does not make them), recommend (the agent proposes actions for human review), draft (the agent prepares complete action proposals but humans approve before execution), human-approved action (the agent executes actions within tight constraints after explicit human approval), and bounded automatic action (the agent acts autonomously within a defined, proven scope).

Each stage requires evidence before the next stage is accessible. A buyer who wants to move an agent from "recommend" to "human-approved action" needs a record of how its recommendations performed: accuracy rate, decision quality relative to human equivalents, failure cases and their resolution. The evidence requirement is not a bureaucratic hoop. It is what keeps the expansion of agent authority tied to demonstrated capability rather than to vendor promises or optimism.

This model also provides a natural structure for contracting agent authority when performance drops. If an agent that was operating at "bounded automatic action" starts producing results that diverge from mandate intent, the response is not to switch it off entirely, it is to step it back to "human-approved action" until the source of the divergence is understood. The model works in both directions.

What the approval threshold should be set to This is the question buyers most frequently ask, and it does not have a universal answer. It depends on campaign scale, the buyer's tolerance for error, the quality of the deal record infrastructure, and the maturity of the inventory relationships involved.

Some useful calibration points. For a campaign with a daily budget under ten thousand pounds, automated execution against a pre-approved mandate is generally appropriate without per-deal human approval, provided the mandate is specific and the deal record is auditable. For a deal above a material threshold (the specific number depends on the buyer's budget scale, but a starting point is any individual deal that represents more than 5% of total campaign budget), a pre-approval step is proportionate. For any deal that modifies the underlying mandate or introduces a new audience signal, pre-approval is necessary regardless of deal value.

These calibrations should be documented in the mandate and reviewed after every campaign. Thresholds that made sense during a pilot will not necessarily make sense when the programme runs at ten times the scale.

The audit question that determines whether any threshold is meaningful Setting a threshold is only useful if the decision log shows what the agent decided, which threshold conditions it evaluated, and whether any decision was escalated for approval. Without that log, a threshold is a rule that may or may not have been followed, and there is no way to tell.

This is the distinction between an agent operating within a threshold architecture and an agent that is described as operating within a threshold architecture. The distinction shows up in the audit. Can a buyer, after the fact, see every decision the agent made, the mandate conditions it checked, and the approval state of each decision that required human sign-off? If yes, the threshold architecture is functional. If no, the appearance of control is not matched by the reality of it.

Buyers evaluating agentic advertising platforms should treat this as a core due diligence question. What does the decision log contain? Who has access to it? How long is it retained? Is it exportable to the buyer's own systems for independent audit?

Where Alkimi's infrastructure sits in this picture The threshold architecture described above requires two things the standard programmatic stack was not built to provide: a shared deal record that captures approval states, and a mandate enforcement layer that sits at the record level rather than inside the agent.

Alkimi's DealSheet is the shared deal record for agent-negotiated transactions in its marketplace. Approval thresholds are enforced at the deal record level: a deal that exceeds the buyer's defined threshold cannot be confirmed without the required approval state being present in the record. The DealSheet is bilaterally held, meaning both buyer and seller reference the same record, so the audit log reflects what was actually agreed rather than what one side recorded.

The earned autonomy model is the governing framework for how Alkimi extends and contracts agent authority within buyer programmes. It is not a marketing description. It is the operational model that determines which decisions get escalated and what evidence is required before scope expands.

Entering Alkimi Marketplace...