
Share:








Share:




Share:




Everyone is doing some version of the same exercise right now. We are putting AI into workflows, automating work that used to sit with people, and sooner or later somebody asks the question that decides whether the whole thing keeps moving: is this actually worth it?
On paper, the economics can look almost absurdly good. A model call costs cents. A task that used to take a person an hour can run in minutes. The spreadsheet practically sells the project for you.
Then production gets involved and the number starts moving. Context gets loaded again. A run repeats. A browser or business system adds its own cost. The awkward case still lands on a person. Different teams start looking at different versions of what the workflow costs.
We have been on both sides of that conversation, building the workflow and then explaining why the savings on the slide are not showing up exactly where expected. We have built the spreadsheet, then watched production rewrite it.
That scar changed how we do the math. We start with the unit of work the operation already understands: one support case closed, one vendor review completed, one incident resolved, one onboarding case completed, one QA cycle closed. Then we ask what it cost to get that unit all the way through successfully.
The model underneath can be GPT, Claude, Gemini, or whatever gets released next month. The unit cost should survive the swap.
And small misses stop being small at volume. In the worked example below, a $0.60 human exception line becomes $15,000 a month at 25,000 cases.
| The number we care about is cost per successful unit of work. Everything else feeds that number. |
The model invoice covers the AI call. The operation pays for more.

Those 4 costs rarely live in the same place. That is where the ambiguity comes from.
Use the unit your operation already counts, then assign every cost in the workflow to that unit. The math gets much easier once everybody is talking about the same denominator.
Take a billing support case almost anyone has seen. A customer writes in because the same purchase appears twice.
The workflow reads the message and account history, checks recent charges, checks the refund rules, and either resolves the case under those rules or sends the exception to a person.
To make the model line concrete, give that case 8,000 input tokens and 1,000 output tokens. At current standard rates on September 24, 2026, both GPT 6 Sol and Claude Sonnet 5 price that token volume at $0.026. We keep both names here on purpose. Swap the provider and only this line of the worksheet should change.
Now add the operating reality. If 10% of cases need 6 minutes from a person at an internal cost of $60 an hour, the expected human exception cost is $0.60 per case.

At 25,000 cases a month, the model line is about $650. The expected human exception line is $15,000. Together they are already $15,650 before any paid tool or retry cost.
The model is around 4% of that combined number.
This is why a cheap model call can produce a bad budget. The question is not whether $0.026 is cheap. It is whether you built the business case around $0.026 while the workflow was quietly carrying another $0.60.
| At this volume, cutting the model price in half would save about $325 a month. Cutting human exception time in half would save $7,500. That tells you where to look first. |
Once the unit cost is visible, the next question is where it is coming from.

These 5 numbers give us enough signal to decide what to inspect next without turning cost management into a finance project.
They also stop a common mistake: cutting the easiest visible line while the expensive part sits somewhere else.
The same monthly spend can come from very different problems. The fix depends on which number is moving.

A high retry rate deserves a reliability investigation. High human minutes point to the exception path. High model and context spend points to what gets sent through the model and which model handles each step.
This is where cost stops being reporting and becomes an architecture decision.
A few rules have survived contact with production for us.
The model catalog will change. These rules survive the release cycle.
This is one reason Mustang sits above the model layer.
Mustang can run on one model or several, from any provider. The Grip keeps company knowledge outside the model. Kernel loads the knowledge a task needs, routes work to a cleared worker, checks the result, and tracks cost per run, execution success, and output quality.
Our current FinOps implementation also reads the provider’s own reported usage instead of guessing it. A spend ceiling can hold a run for a named person rather than letting the cost surface later in a monthly report.
Mustang is currently measured at 35%+ lower cost to run AI in production. We care about that number because production is where context, checks, retries, tools, and people all show up.

Pick one workflow and use a recent batch of real runs.
Add machine spend. Add human minutes. Count retries and failed runs. Divide the total by successful outcomes.
Then circle the biggest line. That is your first optimization target.
That is where we would start too. Give us one workflow, a recent batch of runs, and the systems it touches. We will help you put a real unit cost around it, find where the money is going, and decide which architecture change is worth making first.
| Bring us one workflow where the model bill looks fine and the operation still feels expensive. That is the conversation worth having. |
Share:










We’ve helped teams ship smarter in AI, DevOps, product, and more. Let’s talk.
Actionable insights across AI, DevOps, Product, Security & more