The Cost Lie of AI — What an Agent Really Costs

Above the waterline, everyone argues about the price of fuel. The ship goes down on the iceberg beneath it.


There are companies that now measure their people by token consumption.

I didn’t make this up. There are teams where the monthly AI report carries a new metric right next to revenue and utilization: tokens consumed per head. Use too many, and you get an email.

That’s like judging the economics of a trucking fleet by the price per gallon of diesel — while ignoring what actually determines whether the fleet makes money: full load capacity, on-time delivery, efficient routing.

Fuel is real. It’s on every invoice. But it doesn’t decide whether the fleet earns or burns.

AI works exactly the same way. Tokens are the fuel price. Visible, measurable, on every bill — and more expensive than many would like. But even when it’s high, it doesn’t decide victory or defeat.

Measuring AI by tokens confuses the fuel price with logistics performance.

And because those three metrics — load capacity, on-time delivery, routing efficiency — are precisely the ones your ops and supply chain people already think in, hold on to them. We’ll pick them back up at the end. Each has a counterpart in the AI world that decides success or failure.


The Iceberg

Tokens are the tip above the water. Beneath it lie two layers almost no one accounts for — and the second one sinks projects.

Layer 1 — Visible: The tokens.

Yes, they’re real. And yes, in large deployments they add up to serious numbers — six figures over three years is not unusual. But here’s the point: even if you book the full token cost, it’s only about a fifth of the bill. The bulk lies elsewhere — and no one talks about it.

Optimizing here means haggling over the fuel price while the engine, gearbox, and axles aren’t even paid for yet.

Layer 2 — Below the waterline: The chassis.

This is where the money is. Integration into SAP, MES, the existing system landscape. Guardrails — the five layers that keep an agent from triggering the wrong thing when in doubt (I’ve written about that separately). Monitoring that doesn’t just track uptime but decision quality. Data maintenance that never stops.

This is the running gear. It’s expensive, it’s unglamorous, and it’s what makes the difference between a showroom model and a road-legal truck.


The Honest Calculation

Let’s put numbers on it. A single productive agent in continuous operation, three-year total cost of ownership:

Tabelle kopieren

Item3-Year Cost
Tokens (model usage)€120,000
Integration (SAP, MES, interfaces)€140,000
Guardrails (design, implementation, testing)€90,000
Monitoring & operations (incl. staff)€130,000
Data maintenance (ongoing)€95,000
Total€575,000

The token line accounts for roughly 21 % of the quantifiable total cost.

Now suppose you think that figure is even too high — fine. Even if you double it, you’re adjusting one-fifth of the bill while the other four-fifths stay untouched. That’s the whole point: the line the industry argues about, and the one companies hold their people accountable for, is not the line that decides success or failure.

And that’s just the quantifiable cost. There’s a second bill that never shows up in any spreadsheet.


The Invisible Cost: Deskilling

When an agent makes purchasing decisions for two years, something happens that no invoice records: your people forget how to make them themselves.

This is the most expensive line item of all — and it appears in no budget. A buyer who has spent two years merely confirming what the system proposes loses the judgment that made them valuable in the first place. And with it goes the ability to notice when the agent is wrong.

Which, statistically, it will be. An agent operating at 95 % accuracy per step reaches only about 36 % reliability over twenty sequential steps (the math is in my earlier piece). On-time delivery doesn’t mean “usually” — it means “reliably.”


Bringing the Three Threads Back

Remember the three fleet metrics. Here’s their AI counterpart.

Load capacity → utilization of the process, not the model. The question isn’t how many tokens the agent burns. It’s whether it carries a full, value-adding process end to end — or just automates an isolated task while the process around it still runs on manual labor.

On-time delivery → reliability under compounding error. A chain of agents doesn’t add up its error rates — it multiplies them (as I’ve shown). Reliability isn’t “usually.” It’s “verifiably.”

Routing efficiency → decision latency and cost per process. The real lever isn’t how fast the agent completes a task, but what a complete process costs — including the loops, the callbacks, the corrections. Euros per completed procurement cycle, not seconds per task.

Measure those three, and you’re no longer talking about tokens. You’re talking about value creation.


What This Means for Leadership

The cost question of AI is not an IT question. It’s a leadership question.

Delegate it to IT, and you get an answer about tokens — because that’s the line IT sees. Treat it as a leadership task, and you ask the right questions: What does the chassis cost? What does three years of operation cost? And what will it cost us if, five years from now, our people have forgotten how to decide for themselves?

That’s Adult Supervision: not the refusal to use AI, but the grown-up decision about where it holds and where it ruins. Where it accelerates the process — and where it erodes the very competence that holds your company together at its core.

The fuel price was never the problem.

The problem is the driver who has forgotten how to navigate without GPS — and the fleet that only knows the roads it already drove yesterday.


In your AI business case — are you costing the chassis, or still the fuel?

#AI #SupplyChain #DigitalTransformation #AgenticAI #Industry

E-Mail: sven.vollmer@business-quotient.com

Sven Vollmer is “The Industrial Translator.” He bridges the gap between industrial operational reality (SAP, supply chain) and the possibilities of generative AI. His focus is on value-creating applicationsbeyond the hype.

Transparency Note: This article was created with editorial support from AI (Gemini/Claude). The ideas, technical validation, use case selection, and adult supervision were 100% authored by Sven Vollmer.

LinkedIn: www.linkedin.com/in/sven-vollmer-bq

Similar Posts