Model selection

Two ways the obvious choice costs you money.

A configuration is one model per step — not one model for the whole workflow. TokenWake ranks every possible combination by true cost per delivered task, which is API spend plus modeled failure exposure, and recommends the one that matches the downstream impact you declare, with the shortlist behind it. Here is why that is rarely the one anyone would have picked.

Blind spot one

The winner isn't the cheapest.

Pricing pages and dashboards report API spend — the sticker, and only after the tokens are gone. True cost per delivered task includes retries, abandonment, silent failures, and the human review and rework nobody puts on the invoice. Across the 196 configurations of a draft-and-validate flow, those two measures disagreed by 5.25× — in the opposite direction to the one everybody assumes.

SELECTED CONFIG · API SPEND $0.58 PER DELIVERED UNIT $6.60 CHEAPEST BY API · API SPEND $0.008 PER DELIVERED UNIT $34.65 5.25× MORE EXPENSIVE TO DELIVER THE SAME WORK
Bars show true cost per delivered task — API spend plus modeled failure exposure.

Judged on API spend alone, the cheapest option costs about 75× less to run. Judged on what it takes to actually deliver the work, it costs 5.25× more — the configuration that looked cheapest finished last of all 196, the most expensive decision on the board.

Cheapest is the right answer only if a wrong answer costs you almost nothing. For some workflows that is true. Your report says where it stops being true for yours.

Failure exposure is priced at $25 per silent defect — a declared assumption, not a measurement, and one the conclusion survives at anything above $0.50.

Blind spot two

The right configuration is almost impossible to guess.

Of the 196 configurations you could choose, 183 are beaten on both cost and reliability by something else on the board. And of six standard agentic patterns we swept — ReAct, two variants of Reflexion, two of Plan-and-Execute, and Supervisor/worker — every one came in 1.9× to 2.3× above its own optimum. Nothing about those workflows changed to close the gap, only which model ran at which step.

Here is the whole set — all 196 configurations, every one scored. Narrowing that down to your configuration is standard and defensible statistical method, backed by millions of scenario variation runs.

250 500 750 1,000 1,250 $0.00 $0.25 $0.50 $0.75 $1.00 EXPECTED API $ PER DELIVERED UNIT MODELED SILENT FAILURES PER 1,000
Ruled out (183) Efficient frontier (13)
Every dot is one configuration — what it costs per delivered task, against the silent failures it lets through. Which point on the gold line is top ranked depends on the downstream impact you declare.

See how we get there

What you get

Which one is yours?

Your report names the models. Every step of your workflow with a specific model assigned to it, ranked by true cost per delivered task — plus the shortlist sitting just behind the winner, and the thresholds that would change the answer.

The whole thing, on a real run: a full example forecast — an invented customer, a certified engine run.

Send us one workflow

We will find the most efficient configuration
for the downstream impact you declare.

Plus what it will cost and where the money goes —
all from the same run.

$15,000 introductory offer · one workflow.

Only want your existing configuration priced, without the sweep? Write to contact@tierzerosolutions.io and we'll scope it with you.