1 Oct 2026 · 3 min read

Why some agents finish and others don't: what GRPO tells us about agentic behaviour

The Concourse v1 benchmark put eight AI buyers through 96 identical media-buying runs. Sol and Fable completed every single one. Luna and Haiku completed none. Same brief, same sellers, same rules. The gap is not about intelligence in the abstract. It is about how a model was trained to behave under pressure.

That training story runs through GRPO.

What GRPO actually does

Group Relative Policy Optimization is a reinforcement learning technique introduced by DeepSeek for training language models to reason. The core mechanic is deliberately simple: generate a group of responses to the same prompt, score each one, then use the average score as the baseline. Responses that beat the group average get reinforced. Responses that fall below it get suppressed.

No separate critic model. No external reward model. The signal comes from within the group itself.

The effect on model behaviour is not subtle. When a model is trained this way on tasks with clear, objective outcomes (is the maths correct? does the code run? did the portfolio stay within budget?), it learns that longer, more deliberate reasoning paths tend to win. It learns that checking your work is worth the tokens. It learns to persist through multi-step problems rather than collapsing to a plausible-sounding early answer.

What the Concourse data surfaces

A media-buying run requires a model to hold a brief in mind across several rounds of negotiation with multiple sellers, track a running budget, recognise when a deal does not meet the brief, and assemble a portfolio that satisfies all constraints simultaneously. There is no single clever move that gets you there. You have to reason, update, and stay on task.

The completion table from Concourse v1 makes the reliability difference concrete:

Sol: 12 of 12. Fable: 12 of 12. Opus: 10 of 12. Astra: 9 of 12. Sonnet: 9 of 12. Terra: 3 of 12. Luna: 0 of 12. Haiku: 0 of 12.

Sol and Fable are models with significant reinforcement learning applied to agentic reasoning. The models that struggled are predominantly smaller or less trained on sustained multi-step tasks. This is not a coincidence. GRPO-style training directly rewards the behaviour that completing a Concourse run requires: persistent goal pursuit across a workflow where the right answer is not available at step one.

Completion is not the whole story

This is where it gets interesting. Concourse goes a layer deeper than pass/fail: across the 55 completed buys, 95 of the 110 signed contracts promised a viewability floor below 70%. The agents finished. They just did not negotiate well.

GRPO trains models to complete the workflow. But a reward signal that only asks "did you finish?" will produce agents that finish. If you want agents that also secure favourable terms, hold firm on quality floors, and avoid weak contracts, your reward signal has to capture that. The training shapes the behaviour, but only as far as the reward signal reaches.

This is the practical implication for anyone building agentic systems. You can run a benchmark and see which models complete the task reliably. That tells you a great deal. But the deeper question is: what exactly is the model optimising for? A model trained to complete a media buy will do exactly that. Whether it learns to do it well depends on how you defined well when you built the reward function.

What this means for agentic advertising

Alkimi's premise is that agents can trade media on behalf of humans, with humans retaining approval rights at the deal level. The Concourse data is an early view of what that looks like in practice. Some models are structurally ready for agentic media buying. Others are not, yet.

The gap will close, but not uniformly. GRPO and similar techniques will keep improving reasoning reliability. The harder problem is getting the reward signal right: training models not just to finish deals, but to finish deals that a human buyer would actually approve.

That is the benchmark worth building toward.

All articles

Speak to the team