Inference economics
How falling token costs redraw where value sits in the stack, and what we revise when they fall again.
The falling floor
Every six months the price of a token falls by half, sometimes faster. o4's launch cut reasoning-grade inference forty percent overnight.[1] When the engine gets cheap the profit pool does not vanish, it migrates: away from raw model access, toward whoever owns the defensible layer above it.
Where the margin goes
The margin re-forms around workflow: orchestration, evaluation, and the tool loops that agents run in production.[2] Compute-heavy regions with cheap power set the floor, and grid economics[3] quietly become AI economics.
How we track it
The desk maintains a price-per-token series across six providers, revised whenever list prices move. When a change crosses a threshold, dependent pages get flagged for revision, this page first of all.
- OpenAI o4 launch pricing (Apr 2026)
- AFS dispatch: The coming collapse of the generic AI wrapper
- AFS wiki: Indonesia LNG and the grid