An AI agent optimising a media buy does four things a human buyer used to do by hand: it reads the campaign inputs, forms a view about what to change, takes the action, and records what it did. The difference from older automation is the second step. Programmatic software executed decisions a human had already made. An agent makes the decision itself, inside boundaries the human set, and the quality of the campaign now depends on choices no human watched being made.
TL;DR. An agent reads a brief and live performance data, decides what to adjust, acts on it, and logs the action. It optimises across the same levers a human buyer pulls, targeting, budget allocation, bidding, pacing, and inventory selection, but it pulls them continuously rather than in daily reviews. What separates a well-built agent from a risky one is not how autonomous it is, but where the human sign-off sits and whether every decision leaves an auditable record. The useful mental model, and the one the industry is converging on, is an agent as a fast, tireless junior buyer whose autonomy widens only as its audit trail earns trust.
What does an AI agent actually do in a media buy?
It runs a loop. The agent reads the current state of the campaign, compares that state against the goal it was given, decides on an action, executes it, and then reads the new state to see whether the action worked. That loop runs continuously, which is the first real departure from human buying. A media buyer reviews a campaign once or twice a day and makes a batch of changes. An agent can evaluate and adjust many times an hour, because it is not waiting for anyone to open a dashboard.
The inputs to that loop are specific. The agent reads the brief, which defines the objective, the budget, the flight dates, and the constraints. It reads live performance data: spend, delivery, conversions, cost per outcome, pacing against budget. It reads inventory signals about what is available and at what price. On some systems it also reads the terms of deals it can transact against, which is what lets an agent negotiate rather than only bid.
The output is an action, and this is where the word "optimise" earns its meaning or loses it. A good agent's action is a specific, reversible change with a recorded reason: widen this audience because the current one is exhausted, shift budget to this line item because its cost per outcome is lower, pause this placement because its quality signals dropped. A weak agent produces an action with no recorded reason, which optimises the number on the dashboard while leaving nobody able to explain how.
Which levers does an agent optimise, and how?
The levers are the same ones media buyers have always used. What changes is the speed and the continuity.
Targeting is the first. The agent watches which audience segments deliver against the goal and reallocates delivery toward the ones that work, narrowing or widening reach as performance dictates. A human does this in weekly optimisation passes. An agent does it as the data arrives, which means it can catch an underperforming segment before it burns a day of budget.
Budget allocation is the second. Within the total the human set, the agent moves money between line items, placements, or channels based on where each marginal dollar performs best. This is the lever where autonomy carries the most risk, because a budget shift is a real commitment of spend, and it is the lever most platforms gate behind a value threshold above which a human must approve.
Bidding is the third, and it is the most direct inheritance from programmatic. The agent sets and adjusts bids in real time against its outcome target, the same job a bid algorithm did, but now connected to the agent's wider view of the campaign rather than running as an isolated rule. Pacing is the fourth: the agent keeps delivery on track against the flight, speeding up or slowing down to avoid the two classic failures of underdelivery and front-loaded spend.
Inventory and supply path selection is the fifth, and the one where agents can do something genuinely new. Rather than accepting whatever the pipeline serves, an agent can favour paths with better fee visibility and quality signals, which is where agentic buying intersects with supply path optimisation. The agent is not just choosing what to buy. It is choosing how to buy it.
How AI Agents Optimise Media Buying: A Practical Breakdown
An AI agent optimising a media buy does four things a human buyer used to do by hand: it reads the campaign inputs, forms a view about what to change, takes the action, and records what it did. The difference from older automation is the second step. Programmatic software executed decisions a human had already made. An agent makes the decision itself, inside boundaries the human set, and the quality of the campaign now depends on choices no human watched being made.
TL;DR. An agent reads a brief and live performance data, decides what to adjust, acts on it, and logs the action. It optimises across the same levers a human buyer pulls, targeting, budget allocation, bidding, pacing, and inventory selection, but it pulls them continuously rather than in daily reviews. What separates a well-built agent from a risky one is not how autonomous it is, but where the human sign-off sits and whether every decision leaves an auditable record. Gartner forecasts that 40% of enterprises will demote or decommission autonomous AI agents because governance gaps are identified only after a production incident. The useful mental model is an agent as a fast, tireless junior buyer whose autonomy widens only as its audit trail earns trust.
What does an AI agent actually do in a media buy?
It runs a loop. The agent reads the current state of the campaign, compares that state against the goal it was given, decides on an action, executes it, and then reads the new state to see whether the action worked. That loop runs continuously, which is the first real departure from human buying. A media buyer reviews a campaign once or twice a day and makes a batch of changes. An agent can evaluate and adjust many times an hour, because it is not waiting for anyone to open a dashboard.
The inputs to that loop are specific. The agent reads the brief, which defines the objective, the budget, the flight dates, and the constraints. It reads live performance data: spend, delivery, conversions, cost per outcome, pacing against budget. It reads inventory signals about what is available and at what price. On some systems it also reads the terms of deals it can transact against, which is what lets an agent negotiate rather than only bid. The Ad Context Protocol, the open standard for agent-to-agent advertising maintained by AgenticAdvertising.org, defines that layer explicitly: agents discover inventory, buy media, activate audiences and manage accounts through a common set of tasks rather than through each platform's own interface.
The output is an action, and this is where the word "optimise" earns its meaning or loses it. A good agent's action is a specific, reversible change with a recorded reason: widen this audience because the current one is exhausted, shift budget to this line item because its cost per outcome is lower, pause this placement because its quality signals dropped. A weak agent produces an action with no recorded reason, which optimises the number on the dashboard while leaving nobody able to explain how.
Which levers does an agent optimise, and how?
The levers are the same ones media buyers have always used. What changes is the speed and the continuity.
Targeting is the first. The agent watches which audience segments deliver against the goal and reallocates delivery toward the ones that work, narrowing or widening reach as performance dictates. A human does this in weekly optimisation passes. An agent does it as the data arrives, which means it can catch an underperforming segment before it burns a day of budget.
Budget allocation is the second. Within the total the human set, the agent moves money between line items, placements, or channels based on where each marginal dollar performs best. This is the lever where autonomy carries the most risk, because a budget shift is a real commitment of spend, and it is the lever most platforms gate behind a value threshold above which a human must approve.
Bidding is the third, and it is the most direct inheritance from programmatic. The agent sets and adjusts bids in real time against its outcome target, the same job a bid algorithm did, but now connected to the agent's wider view of the campaign rather than running as an isolated rule. Pacing is the fourth: the agent keeps delivery on track against the flight, speeding up or slowing down to avoid the two classic failures of underdelivery and front-loaded spend.
Inventory and supply path selection is the fifth, and the one where agents can do something genuinely new. Rather than accepting whatever the pipeline serves, an agent can favour paths with better fee visibility and quality signals, which is where agentic buying intersects with supply path optimisation. PubMatic has built this into its agentic stack as a Fee Transparency Agent and a Detailed Reasoning Agent, aimed at accounting for what each dollar in the supply chain is doing and why the system chose as it did. The agent is not just choosing what to buy. It is choosing how to buy it.
What decisions does the agent make on its own, and what needs a human?
This is the question that separates a defensible deployment from a reckless one, and the honest answer is that it depends entirely on how the boundaries are drawn. The technology does not decide this. The buyer does, when they configure the agent.
The pattern the industry is settling on is a graded scale of autonomy rather than an on-off switch, and the analyst case for it is explicit. Gartner's Shiva Varma has argued that enterprises treating agent governance as binary, "either locked down or fully trusted", is itself the root cause of failure, because uniform controls either over-restrict simple agents or under-restrict autonomous ones. In Gartner's framing, the highest autonomy level is one where agents act independently inside defined guardrails while humans review exceptions, audit logs and aggregated outcomes rather than individual decisions. Applied to media buying, that means low-stakes reversible actions, a small bid adjustment or a pacing tweak, get automated first, while large budget reallocations and the opening of new terms with a seller stay behind an approval gate longer.
The specifications are being written to that shape. The IAB Tech Lab's AAMP framework builds bounded autonomy into the architecture: in the AAMP 2.3 release, any path that commits spend is designed to be deterministic and provable, with human approvals required outside set value thresholds, and agent negotiation is constrained by deterministic guardrails rather than left open-ended. Platform implementations take a similar shape. PubMatic's guardrail architecture for AgenticOS runs a five-step governance framework in which agents inherit business rules, draw only from pre-approved inventory, creative and audience pools, and escalate through authenticated approval workflows. The common thread is that autonomy is bounded by design, not granted wholesale.
The reason this matters is a failure mode specific to autonomous systems. A human buyer who misreads the data makes one bad change and notices at the next review. An agent that misreads the data makes the same bad change repeatedly, at speed, because it is doing exactly what it was told to do against a flawed signal. Varma's description of the problem is that autonomous actions execute at a scale and speed capable of outpacing human oversight while accountability for the outcome stays with the organisation. Both of the men writing the competing agentic protocols named the same underlying risk independently, in written answers to the trade publication ADOTAT: hallucination, treated not as a curiosity but as the thing the architecture has to be engineered to survive. The human approval gate is not bureaucratic caution. It is the interrupt.
How is this different from the automated bidding buyers already use?
Automated bidding optimised one variable inside fixed rules. You set a target cost per acquisition and the algorithm adjusted bids to hit it, but it could not decide to move budget, change the audience, or question the inventory. It was a specialist doing one job very fast within lines a human drew.
An agent optimises across all the levers at once and, more importantly, reasons about the relationships between them. It can recognise that the reason cost per outcome is rising is not the bid but the audience, and act on the audience instead. That connected judgement is the capability that is new. It is also the capability that creates the auditability problem, because a decision that draws on several signals and a chain of reasoning is harder to reconstruct after the fact than a rule that either fired or did not. Forrester makes the technical version of this point: agent reasoning involves branching logic, internal state transitions and multistep transformations, and conventional logs cannot trace a reasoning chain, which makes forensics and root-cause analysis materially harder than they were.
So the practical upgrade is not speed. Programmatic already delivered speed. The upgrade is that the machine now handles the judgement between the levers, and the practical cost is that the record of that judgement has to be captured well enough that a human can inspect it later. An agent that optimises brilliantly but records nothing has moved the buyer's problem, not solved it.
What should a buyer put in place before letting an agent optimise?
Treat the agent as a new hire with unusual speed and no common sense about consequences, and manage it accordingly. Three things need to be in place before it touches live budget.
First, defined boundaries. What can the agent change without asking, and what requires sign-off? Draw the line by reversibility and by spend at risk, and start it conservative. It is easy to widen an agent's autonomy once it has earned trust and painful to claw it back after it has spent money you cannot recover.
Second, a real audit trail. Every action the agent takes should be recorded with the reasoning behind it, in a form a human can read months later. The standards bodies now say this in the open. Shailley Singh, managing director for product at the IAB Tech Lab, told ADOTAT that "an agentic transaction cannot simply rely on a model saying that a transaction occurred", and that the architecture needs a deterministic, auditable record of what was proposed, approved and executed, with explicit transaction states both parties can reference and check rather than an agent confirming a buy in conversational English. The test for a buyer is narrower and more practical: whether someone who was not there can reconstruct why the agent moved budget in week three. Ask also who can verify the record besides the vendor that wrote it. Reviewing PubMatic's governance launch, PPC Land noted that the announcement described no third-party attestation, no exportable log format and no independent standard to check the audit trail against, which leaves the difference between a vendor-published record and an independently verifiable one unresolved. That gap is the one to press on in a pilot conversation.
Third, a kill switch and a proof period. The agent should run in a bounded pilot, recommend-only or tightly capped, long enough to demonstrate that its decisions hold up before it graduates to wider control. Forrester's guidance is to scale in stages, starting with bounded tasks behind approval gates and rollback paths, and to "widen autonomy only when the controls earn it". A pilot is also closer to the current state of the category than most keynotes suggest. Asked in writing whether a binding, signed record of an agentic transaction exists and works in production today, Brian O'Kelley, founder of the Ad Context Protocol, answered that it exists and works but is not yet active in production at scale. Andrew Mole, chief executive of the publisher platform pubX and a founding member of the Agentic Advertising Organization, is among the few operators running live agent-bought volume, and put his own figure at two to three thousand dollars a day of gross media spend across a handful of publishers, and not continuous. Mole's company sells the thesis that the existing stack can be skipped, so his forecasting carries a commercial interest, but he is one of the only people in the category to answer the scale question with a number rather than a roadmap. Pilot-stage autonomy should never quietly become full budget control without proof, and the fact that almost nobody is running this at volume yet is the argument for building the record now rather than retrofitting it later.
An agent optimising a media buy is not magic and it is not a threat. It is a fast, literal, tireless worker that will do exactly what its boundaries allow, well or badly, and leave you to answer for the result. The buyers who get value from it are the ones who spend as much effort on the boundaries and the record as they do on the goal.
This article references the IAB Tech Lab's published AAMP specifications, the Ad Context Protocol maintained by AgenticAdvertising.org, PubMatic's AgenticOS and guardrail architecture, Gartner and Forrester research on agentic AI governance, and on-record written answers from Brian O'Kelley, Shailley Singh and Andrew Mole obtained and published by ADOTAT. Where it describes autonomy and approval models, it reflects the graded-autonomy approach documented across those sources rather than the design of any single platform.