The Real Cost of AI Agents Is Not Tokens
McKinsey's latest analysis of agentic AI economics shows why token prices tell only part of the story. The real question is what it costs to complete the work.
McKinsey published an interesting analysis this week about the economics of agentic AI, and one number immediately stood out.
In one of its banking examples, large language model tokens account for only 20 to 25 percent of the variable cost of running an AI agent. Human oversight, on the other hand, accounts for 70 to 75 percent.
That changes the conversation.
For the past few years, much of the discussion around the cost of AI has focused on tokens. How many tokens does a model consume? Which model is cheaper? How much does one million input or output tokens cost?
Those questions matter, but once AI moves from assistants that answer questions to agents that perform actual business processes, they are no longer the most important questions.
The better question is: what does it cost to get the work done?
Cheap tokens do not automatically mean cheap AI
There is an interesting contradiction happening in AI. Model prices continue to fall, while enterprise spending on AI continues to increase.
McKinsey describes the emerging discipline of managing these costs as "tokenomics" or AI economics. The underlying problem is relatively easy to understand when you look at what an agentic workflow actually does.
An AI assistant might receive a prompt and produce an answer. An agentic workflow can go through many more steps before the work is finished. It might:
- classify an incoming request and determine its intent
- extract relevant information
- retrieve knowledge or historical context
- query a CRM, ERP or another business system
- reason about the information it finds
- call APIs or trigger other actions
- generate and validate an answer
- involve a human when necessary
One business transaction can therefore trigger many different AI and software operations.
The cost of the individual token is only one small piece of a much larger economic puzzle.
You do not need a Ferrari to go to the bakery
We have used this analogy internally at ReplyFabric for quite some time: you do not need a Ferrari to go to the bakery around the corner.
And you certainly do not need the most powerful reasoning model available to detect whether an email is written in Dutch or English.
McKinsey makes essentially the same economic argument. Model routing and rightsizing are becoming important parts of AI economics, alongside caching, reuse, batching and reducing unnecessary inference.
A well-designed agentic system therefore should not simply ask which AI model is best. It should determine which model, system or technique is best suited to each individual task.
Sometimes that means using a frontier model because the task requires sophisticated reasoning. For another step, a lightweight model may be more than sufficient. And sometimes the correct solution is not an AI model at all, but a deterministic rule, database query or API call.
This distinction becomes enormously important once a system starts processing thousands or millions of transactions.
But model costs are only part of the story
This is where McKinsey's analysis becomes particularly interesting.
In its illustrative banking example, the distribution of variable costs for an AI agent looks approximately like this:
- 70 to 75 percent for human oversight
- 20 to 25 percent for large language model tokens
- 3 to 4 percent for AI infrastructure
- 2 to 3 percent for APIs and add-ons
SOURCE: McKinsey, August 2026
¹Customer-facing agent that drives activities with direct profit-and-loss impact. ²Midpoint values used. ³10–20% of runs reviewed by human full-time employees (FTEs) for ~5 minutes each, at ~$100,000 annual cost per FTE. ⁴Approximately 70% of runs use frontier or higher-cost models at published token pricing; each run averages 12 turns, 4,000 input words, and 500 output words per turn, with ~1.3 words per token. ⁵Orchestration, model gateway, observability, and vector database reads priced per turn, and vector database writes priced per run, using Langfuse, LangGraph, and Portkey benchmarks. ⁶Approximately 10–15 tool and data API calls per run. ⁷Approximately 500 agents per enterprise, managed by agent operations FTEs at ~$150,000 per year and 20 agents per FTE. ⁸Approximately $80,000 per year platform cost based on Azure benchmarks, covering a 500-agent fleet with no corporate discount.
Companies can spend a lot of energy trying to reduce the price of model calls while overlooking the much larger cost of human intervention. McKinsey therefore argues that organizations should look beyond model selection and token optimization. Redesigning the workflow to reduce exceptions and simplify expensive human review can have a much larger economic impact.
That does not mean removing humans from the process. It means designing a better human-in-the-loop process.
There is a huge difference between asking an employee to verify a correct proposal in 20 seconds and asking the same employee to spend five minutes correcting an unreliable AI response. Both systems technically have a human in the loop, but their economics are completely different.
Accuracy becomes an economic metric
We often discuss AI accuracy as a technical or quality metric, but in operational AI it has a direct economic consequence.
Imagine one AI workflow costs €0.03 per transaction but frequently requires several minutes of human correction. Another costs €0.10 but produces an output that can usually be reviewed and approved within seconds.
The second workflow can easily be cheaper overall.
Optimizing AI economics is therefore not synonymous with choosing the cheapest model. The objective should be to minimize the fully loaded cost of completing the work while maintaining the required quality, control and compliance.
This also changes how we should look at familiar AI metrics. Accuracy, exception rates, human review time and acceptance rates are not merely technical KPIs. They contribute directly to the economics of the process.
And that leads to what I think is the most important concept in McKinsey's article.
Completed work ROI is what really matters
McKinsey calls it completed work ROI.
Instead of calculating what one agent costs, calculate what it costs to finish the actual business process. That could be a completed customer onboarding, a resolved insurance claim, a closed sale or any other meaningful outcome.
The calculation should include everything involved:
- AI models and agents
- deterministic software and business rules
- APIs and external services
- infrastructure and orchestration
- human review and intervention
- exception handling
McKinsey illustrates this with customer onboarding in banking. What looks like one relatively straightforward process can actually require five to seven specialized agents, more than three deterministic rule engines and two to four human teams providing oversight.
Despite all that complexity, McKinsey estimates that agentic AI could reduce the total cost of completing the onboarding workflow from approximately $50 to $150 per customer to around $10 to $30.
That's the metric that matters.
Not the number of tokens consumed or prompts executed, and not even the number of agents deployed. What matters is the work that gets completed and the total cost required to complete it.
High volume can actually make AI more attractive
There is another counterintuitive lesson in the report. Companies sometimes worry that high-volume AI processes will automatically become prohibitively expensive, but McKinsey shows why the opposite can happen.
Agentic systems have fixed costs. Infrastructure, orchestration, monitoring and development need to exist regardless of whether an agent performs 100 or 10,000 tasks. As volume increases, those costs can be distributed across more transactions.
McKinsey's illustrative numbers make the effect very clear:
- At 100 runs, the cost is around $78 to $79 per run
- At 2,500 runs, it falls to around $4 to $5
- At 5,000 runs, it falls further to around $3 to $4
- At 10,000 runs, the cost is around $2 to $3
The second economic advantage is reuse. McKinsey recommends building capabilities once and reusing them across multiple high-value workflows instead of creating isolated agents for every individual problem.
Together, these observations suggest a useful rule when looking for good agentic AI opportunities: look for large pools of valuable, repeatable work where the same capabilities can be reused.
A spectacular AI demo that runs 50 times per year may have much less economic value than a relatively boring workflow that happens 100,000 times.
AI can also make previously uneconomic work possible
Cost reduction is only one side of the equation.
One of the more interesting observations in the McKinsey report is that AI can change the business model itself. Human expertise has traditionally been scarce and expensive, which means companies have often had to reserve highly personalized services for their most valuable customers.
AI changes some of those economics because cognition can increasingly be scaled.
McKinsey gives personalized banking services as an example. Support that historically made economic sense only for high-value customers could potentially be offered to a much larger customer base.
The same principle could apply across many industries. Personalized advice, support, analysis and follow-up that previously required too much human time may suddenly become economically feasible.
That creates two fundamentally different forms of agentic AI ROI:
- Efficiency by performing existing work faster or at a lower total cost
- Expansion by making new services or levels of personalization economically possible
The second category may ultimately prove even more transformative than the first.
Why operational email fits this economic model
This way of looking at AI is one of the reasons we focus ReplyFabric on shared mailboxes.
A shared mailbox is not particularly exciting technology. The work happening inside it, however, can be extremely valuable and highly repetitive.
A reservations mailbox receives reservation requests, cancellations, changes and questions every day. A finance mailbox handles invoices, payment questions and document requests. A customer service mailbox may receive thousands of recurring questions and cases, while an operations mailbox receives requests that need to be understood, enriched with business information, assigned and answered.
The individual emails are different, but the underlying operational patterns repeat continuously.
That combination of volume, repetition and business value is exactly where agentic economics become interesting.
It also allows the same underlying capabilities to be reused. Language detection, categorization, intent recognition, feature extraction, knowledge retrieval, business data lookups, routing, reply generation and validation can support many different workflows rather than being built from scratch for every use case.
We deliberately do not sell tokens
There is another part of McKinsey's analysis that resonates strongly with how we designed ReplyFabric's pricing.
AI economics are changing continuously. Models change, prices change, capabilities improve, providers introduce new models and workflows evolve. McKinsey argues that organizations need to continuously manage those changing economics.
We agree.
We just do not think our customers should have to manage that complexity themselves.
ReplyFabric therefore does not charge customers based on tokens, prompts, inference calls or the particular AI model being used. We price something businesses already understand: emails processed.
Our plans include 2,500, 7,500 or 15,000 processed emails, with additional processing priced at €0.08 per email.
An email costs the same to process regardless of what happens underneath. Depending on the workflow, ReplyFabric may need to:
- categorize the email and determine its intent
- analyze priority or sentiment
- extract structured information
- retrieve knowledge and historical context
- look up information in business systems
- generate a proposed reply
- validate the result before presenting it to the user
Different models and deterministic systems can be involved at different stages. The customer does not need to calculate any of this.
They know their email volume, so they can predict their cost.
We manage the AI economics underneath.
The complexity should be ours, not the customer's
A fixed processing price does not make AI economics disappear. It transfers responsibility for managing those economics from the customer to ReplyFabric.
And we think that is exactly where that responsibility belongs.
We continuously need to determine which model provides the right balance of capability, accuracy, latency and cost. If a lightweight model can reliably perform a particular task, there is no reason to use an expensive reasoning model. If a more capable model substantially improves quality for another task, paying more for that individual step may make perfect economic sense.
As models become cheaper or more capable, we can take advantage of those improvements without asking customers to rethink their budgets or understand what happens behind the scenes.
You should not need a degree in token economics to predict next month's software bill.
Human-in-the-loop does not mean human-does-the-work
McKinsey's finding about the cost of human oversight also reinforces another principle we care about at ReplyFabric.
Keeping humans in control is important, particularly when AI is communicating with customers or handling operational business processes. But human control should not mean rebuilding the work the AI just performed.
The goal is to make human intervention efficient.
A good human-in-the-loop workflow allows someone to read, check, adjust when necessary and approve. As quality improves, fewer cases should require substantial intervention and more should be handled with quick oversight.
For ReplyFabric, that means metrics such as draft acceptance and editing behavior are not merely product-quality statistics. They are also indicators of economic value.
If an employee previously needed five minutes to investigate and answer an email and now needs 30 seconds to verify and approve a prepared response, that difference matters far more economically than shaving a fraction of a cent from the underlying model call.
From FinOps to AgentOps
McKinsey makes another interesting comparison with cloud computing.
As cloud adoption grew, organizations developed FinOps to actively manage cloud consumption and spending. McKinsey expects a similar discipline to emerge around agents and calls it AgentOps.
Organizations will increasingly need to understand:
- which workflows are generating measurable value
- which agents or capabilities should be reused or consolidated
- where cheaper models can deliver the same quality
- where expensive human exceptions are hurting the economics
- which workflows should be redesigned or even discontinued
McKinsey goes as far as recommending that organizations include AI unit economics in their quarterly business reviews.
That tells us something important about where enterprise AI is heading. AI is moving out of the experimental budget and becoming operational infrastructure. Once that happens, companies need to manage not only what AI can do, but whether the economics continue to make sense.
The real AI KPI is not how much AI you use
For a while, companies could demonstrate AI progress by showing how many employees had access to an AI assistant, how many prompts were being generated or how many AI pilots had been launched.
That phase is ending.
The more useful questions now are about outcomes. What work did the AI actually complete? How much human effort remained? How reliable was the result? What did the complete workflow cost, and what economic value did it create?
McKinsey's report makes a convincing case that this is where the conversation around agentic AI needs to go.
For AI providers, I think there is another lesson as well. Customers should not need to understand the economics of every token, model and inference underneath an AI product. They should be able to understand the economics in terms of their own business.
For ReplyFabric, we deliberately make that unit simple:
One email in. One predictable processing price. However complicated the AI underneath needs to be.
Frequently Asked Questions

About the Author
Tom Vanderbauwhede is the founder & CEO of ReplyFabric, lecturer in AI at KdG University, and a seasoned entrepreneur with 25+ years of business experience. He holds master's degrees in Applied Economics, Business Administration (MBA), and Strategic Change Management & Leadership. Tom is passionate about building AI tools that reduce email overload and help teams focus on what matters.
Connect with Tom on LinkedIn and follow his journey as a founder.