Agent pilot cost and reliability model
A calculator for whether an agentic AI pilot is worth building, once you count the tokens a failed run still burns and the human review time it still requires.
What it shows
- Pricing an AI pilot on completed outcomes, not attempted runs
- Making retry cost and residual human review time visible
- Finding the volume where a pilot actually breaks even
A demonstration built from public and synthetic data, not from any real client or organization.
The cost of a completed, acceptable outcome, and the human time the pilot still consumes after it ships, are the two numbers that decide whether an agent pilot is worth building. Both run higher than the per-run token estimate a pilot usually gets approved and renewed on.
This model asks for the inputs a team already knows: how many runs a month, how many steps a run, roughly how many tokens a step, what share of runs a person still reads, and how long that reading takes. It then separates cost per attempt from cost per success, because a run that fails still spends tokens and still spends reviewer minutes without producing anything.
The starting numbers are sourced: token prices match current published API pricing and the hourly rate matches the U.S. Bureau of Labor Statistics’ latest measure of loaded labor cost, both cited on the page. Enter your own numbers if your situation differs.
A demonstration. The starting numbers are cited (see sources above), but every field is yours to edit. This is a planning aid, not a forecast, and it does not provide financial advice.
Bring us the decision, process, or system that is not working.
We will help you understand the problem, determine what the evidence supports, and build a better way forward.