Worked scenario
The April Problem
In 2026 a large engineering organization exhausted its annual AI budget in four months, then capped what its engineers could spend. That much is public record.
Nobody published the inside of it, so we did not model their company. We modeled the pattern — a coding assistant, deployed at scale, running on its own. Every input that takes is knowable in January: headcount, the adoption plan, how a session retries, how finished work spawns more.
The reasoning
Every number here comes from something we declared.
We publish this to show the reasoning, not to prove we guessed a number. A forecast is only as good as its inputs, so here are ours — separated into what is public and what we chose.
- Public — roughly 5,000 engineers, and the fact that a spending limit was imposed after four months.
- Declared by us — session volumes, retry behaviour, and how finished work spawns more work. This is the structure the forecast is built from.
- The budget line is ours. Nobody published theirs, so we drew an illustrative one. The dates below are where our line crosses — not when their money ran out.
- The cap is a proxy for the reported one, and the alternative models in the third scenario are an assumption — not a quote, not a recommendation.
- Zero silent failures is assumed throughout. A real forecast prices them, which makes every number below conservative.
The forecast
A budget that was gone before spring.
Three ways the same four months could have run. The work is identical in all three — same volume, same behaviour.
Modeled cumulative spend, January through April, against the illustrative budget line.
- No effective limit — $2.206M spent, crosses the line on April 6. No work stopped.
- Cap what each engineer can spend — $1.582M, crosses on April 29, and 20,459 sessions stop mid-work.
- Keep the cap, change which models do the work — $689K, never crosses, and a third as much work is lost.
The finding
Three findings — only one solution — TokenWake.
1 · The spend was running away.
Left alone, spend reaches $2.206M by April 6, on a budget built to last the year. Session cost runs about 74× from the middle of the range to the top percentile, so the per-seat average anyone would budget from cannot see it coming.
TokenWake would not have forecast this workflow. A flow with no effective ceiling gets named and refused, not priced — every scenario here is bounded, and the first one's ceiling is simply set high enough never to matter.
2 · Capped, the budget still ran out.
$623K
saved by the cap — 28.3% of spend
23 days
longer before it ran out
20,459
sessions stopped mid-work
TokenWake tells you where the work stops — which step a task halts at under a given limit, and which later steps stop with it — before you pick the number.
3 · Keep the cap, assign the right models.
$689K over the same four months, with the same cap still in place. It never crosses the line at all, and it stops a third as much work. Same behaviour, same volume — different models doing it.
The right model for each step is what TokenWake computes. Not for this study, though: the configuration behind Scenario 3 was assumed, against a company nobody has the inside of. Naming the models takes a real workflow — your steps, your volumes, your cost of a wrong answer.
Audit snapshot: July 26, 2026. Methodology brief and assumption register available on request.
What this means for your deployment
Ask it in January, not April.
If you're deploying agents that work on their own,
you are in the January of this story.
Send us the workflow and we'll run it.
$15,000 introductory offer · one workflow.