The trend

The cost per token of running a given quality tier of model has dropped sharply year over year, driven by more efficient accelerator hardware, better serving infrastructure, and smaller distilled models that match older, larger models' quality.

Why it matters commercially

Falling inference costs are what make agentic workflows — which may call a model dozens of times to complete one task — commercially viable at all; the same workflow that was cost-prohibitive two years ago is now routine.

What to watch

As costs fall, competitive pressure shifts from 'can we afford to run this' to 'how much quality can we buy at the same budget' — expect continued pressure toward smaller, more efficient specialist models for well-defined tasks.

Building something in this space? Our studio helps teams design and ship the AI agents behind trends like this one.

Talk to us →