7 Oct 2026 · 5 min read
What your AI buyer is actually trained to optimise for
An AI buyer does what it was trained to do. The Concourse v1 benchmark found that models trained to complete media buys did exactly that: 55 of 96 runs finished with signed portfolios. In 95 of those 110 contracts, the agent accepted a viewability floor below 70%. The training worked. The reward signal stopped too early.
At a glance
Reward signal: the feedback mechanism that tells a model whether a response or action was good. In reinforcement learning, the model learns to maximise the reward. Define the reward poorly and the model optimises for the wrong thing.
Completion bias: the tendency of an agent trained to finish tasks to treat finishing as the goal, regardless of the quality of the outcome. A buyer with completion bias closes deals; it does not necessarily close good deals.
Contract quality: the terms an agent accepted in a completed deal: viewability floors, CPM rates, brand safety commitments. These are independent of whether the task was completed.
Human approval layer: a governance step in which a human reviews and approves each deal an agent proposes before it is activated. The layer that catches what the reward signal missed.
Why training shapes agent behaviour in media buying
AI buyers are language models that have been trained to reason through multi-step tasks. The training process involves running the model on many examples of a task and reinforcing the responses that score well. The score is the reward signal.
When the task is completing a media buy and the reward is based on whether a portfolio was assembled and signed, the model learns to assemble and sign portfolios. It does not learn to negotiate well unless the reward signal explicitly captures negotiation quality.
This is not a flaw. It is how the technology works. The model is doing exactly what it was trained to do. The question is whether what it was trained to do is what you actually need.
What the Concourse data shows about reward signals
Concourse v1 ran eight AI buyers through 96 identical media-buying runs. The models with the highest completion rates, Sol and Fable, finished all 12 of their runs. Models with lower completion rates failed on multiple runs.
Across the 55 completed buys, 95 of 110 signed contracts accepted viewability floors below 70%. Sol and Fable completed every run. They also produced contracts that would not pass a standard human quality review.
The training that makes these models reliable completers does not automatically make them good negotiators. Completion reliability and negotiation quality are different capabilities that require different reward signals to develop.
The difference between completing a buy and doing a good buy
A completed buy means a portfolio was assembled, contracts were signed, and the budget was not exceeded. A good buy means those contracts also met the buyer's quality standards: viewability floors, CPM ceilings, brand safety commitments.
A human media buyer is trained, explicitly or through experience, to care about both. They know that closing a deal with a weak viewability floor is not a win. An AI buyer trained only on whether the deal closed does not have that knowledge unless it was built into the training.
The Concourse data makes this concrete. 86% of completed contracts missed the viewability floor. The agents closed. A human buyer looking at those contracts would not consider them acceptable.
What this means for buyers evaluating AI agents
Ask vendors what the agent was trained to optimise for. A vague answer is informative. If a vendor cannot describe the reward signal their model was trained on, you cannot predict how it will behave under pressure.
Ask for contract-quality data, not just completion-rate data. What viewability floors did the agent accept? What CPM rates? Were any constraint violations flagged? These numbers tell you what the agent actually optimised for.
Build a human approval step into the deal workflow regardless of the agent's training history. An agent proposes; a human approves. This governance layer catches the gap between what the model optimised for and what your buying standards actually require.
Frequently asked questions
What does it mean that an AI buyer is trained to optimise for task completion? It means the model learned, through its training process, that finishing the task is what gets rewarded. If that reward signal did not include contract quality, the model has no reason to prioritise it.
Can an AI buyer be retrained to optimise for contract quality? Yes, but the reward signal must capture quality explicitly: viewability floors met, CPM at or below ceiling, brand safety criteria satisfied. Retraining requires both a clear quality definition and data to train against.
Why did the Concourse benchmark show poor contract quality even from high-completion models? Because the training that produces high completion rates rewards finishing, not quality. A model can be excellent at completing tasks and poor at negotiating terms if those were trained separately.
How does the human approval layer fix the reward signal problem? It does not fix the model's training. It adds a checkpoint before the agent's contracts are activated. A poor-quality contract gets rejected before it costs anything, regardless of what the model was trained to do.
What questions should I ask an AI buyer vendor about training? Ask what the reward signal was during training. Ask for contract-quality data from benchmark or pilot runs, not just completion rates. Ask how the agent behaves when it cannot meet brief constraints.