The Token Furnace
Turns compute, retries, and human cleanup into cost without the work.
The team has found a cheaper model. The price per million tokens is a fraction of what they were paying, the new agent runner is impressively small, and somebody has posted the savings in the engineering channel. A week later, the agent takes more attempts to finish each task. Difficult cases fall back to the original model after the cheaper one has spent a while trying. An engineer checks the output every afternoon and quietly fixes the parts that almost worked. The model line item is down. The work has acquired a second shift.
Symptom
Savings are celebrated per token while total spend climbs. Idle capacity bills around the clock. Cleanup hours never appear on the same dashboard as inference costs.
Why It Matters
The Token Furnace turns compute, retries, and human cleanup into cost without enough useful work to justify it. A cheaper model can produce a more expensive result.
What the Chapter Gives You
How to account for retries, fallbacks, and cleanup in agent economics, the idle-capacity trap in always-on fleets, and the unit that actually measures agent value.
Want the full chapter? Grab the free cheat sheet, read an excerpt, or get the book.
Recognize this one in your codebase?
Free cheat sheet, excerpts, and interactive diagnostics.