When an engineering leader tells me their developers ran out of tokens, I hear good news. Somewhere in the org chart, someone treats that as a budget overrun to investigate. It is the opposite. It is the clearest adoption signal an enterprise gets.
Do the math on what a token cap actually protects. A senior engineer's time costs the company a couple hundred dollars an hour, fully loaded. The tokens required to let an AI agent chew on the same problem for that hour cost a few dollars. When a developer hits a cap and stops delegating work to the model, you have preserved a rounding error in the AI budget by idling the most expensive resource you have. I have watched companies debate a small monthly token upgrade for engineers whose loaded cost runs well into six figures. That debate should take four seconds.
The better question for executives is the ratio question. What fraction of your engineering spend should go to humans and what fraction to tokens? Most enterprises today sit around 99 to 1. The frontier startups building with autonomous agent loops are trending toward something much closer to even, because they discovered that letting an agent churn on a problem for hours costs less than interrupting a human for ten minutes. Your ratio does not need to match theirs. But if it is not moving, your AI program is not real yet.
Per-task costs will keep dropping. Open models are already good enough for a large class of business process work, and that pressure on pricing is permanent. But you should still expect to spend more on tokens next year than this year, because operationalizing AI multiplies use cases faster than unit costs fall. That dynamic is the heart of the token reckoning that is coming for every enterprise budget.
So treat token exhaustion the way you treat a product hitting capacity limits, as demand outrunning supply. Fund it, measure the ratio, and watch which teams keep hitting the ceiling. Those are the teams showing you where the leverage is.
Key takeaways
- Teams that exhaust their token budgets are the teams actually shipping with AI, and that demand signal should be funded, not throttled.
- Per-task token costs are falling, but use cases are multiplying much faster, so total token spend should grow year over year.
- The right question for executives is not how to cap token spend but what ratio of human cost to token cost the organization is aiming for.
FAQ
Why would rising token spend be a good thing for an enterprise?
Because token spend tracks actual usage. A team burning through its token allocation is a team delegating real work to AI. Flat token spend usually means the tools are licensed but idle, which is the far more expensive failure mode.
Won't falling model prices reduce total AI spend over time?
Per-task costs will keep falling as open models pressure frontier pricing. But organizations that operationalize AI find ten to a hundred times more use cases, so aggregate spend rises even as unit economics improve. Efficiency expands the market, it does not shrink it.
Related Essays
The Token Reckoning Is Coming
Engineers burn ten to fifteen million tokens a day. Powerful models run tasks that do not need them. By the end of 2026, the CFO will start asking questions and most teams will not have answers.
You Are Underspending on Tokens, Not Overspending
The instinct heading into 2027 is to optimize token spend. Most engineering teams have the opposite problem, they have not yet reached the spend that quality actually requires.
The Fifth Inning and the First
San Francisco is mid-game on agentic engineering while most enterprises are just starting. That diffusion lag is not a problem. It is the market.
Drowning in pull requests that need your review? Try Tembo Review, a beautiful AI-assisted PR review tool unlike anything you’ve used.