Gartner published a number on August 17, 2026 that should change how you price your product. AI inference cost per agentic workflow will increase more than fivefold through 2028. At the same time, per-token prices keep falling. Both statements are true, and the space between them is where your gross margin goes to die.
Here is the damaging admission up front: the fivefold figure is a forecast, not a measurement. Nobody has audited 2028. But the mechanism behind it is already visible in production bills today, and it is boring, mechanical, and easy to model on a napkin.
Gartner senior director analyst Will Sommer calls it the "inference paradox." Better unit economics are pushing total AI cost up, not down. Cheaper tokens tempt teams to build more complex workflows, and those workflows eat tokens faster than prices drop.
If you sell an agent product on a flat monthly fee, you are quietly short a rising cost curve. Let me show you the math and what to do about it this week.
The Scissor Gap
Picture a pair of scissors. The bottom blade is price per token, falling hard. The top blade is tokens per completed task, rising harder. Your margin is the paper in between.
Four numbers that decide whether your agent product makes money in 2028.
Blade one is real. GPT-4-class pricing has collapsed from roughly $30 per million tokens in 2023 to around $0.10 per million in 2026, according to a Zowie analysis published August 5, 2026. NavyaAI reported in February 2026 that per-token prices have dropped roughly 99.7% since GPT-3-era rates. Vendors love this chart, and they should. It is genuine progress.
Blade two is the one nobody puts in the deck. Splunk reported on June 29, 2026 that agentic systems burn 5 to 30 times more tokens per task than a single chatbot call. Redress Compliance put the agent-loop multiplier at 3 to 10 times a single-shot prompt on the same task, with average input tokens per query up 2 to 4 times versus a 2024 baseline. McKinsey goes further, estimating that some agentic tasks consume around 1,000 times the tokens of a simple chat exchange.
The Scissor Gap, in one sentence: unit prices fall on a smooth curve, task complexity rises on a step function, and the step function wins. Gartner's framing names three drivers. Model economics improve, that improvement unlocks more powerful and more expensive models, and those models get pointed at workflows that consume far more tokens than a chatbot ever did.
Cost Per Successful Workflow Is the Only Price That Matters
Most teams learn this the hard way, and the hard way hurts.
You build a demo. It costs almost nothing. You price a seat at $50 per month because your competitor charges $49. You launch, usage climbs, and six months later finance asks why the model invoice grew 4x while headcount stayed flat. Nobody on the team can answer, because nobody instrumented cost per task.
Now the easy way. Stop tracking cost per token and start tracking cost per successful workflow. Say your agent runs 200 workflows per seat per month and each one costs you $0.08 in inference. That is $16 of cost against $50 of revenue, so 68% gross margin, and you feel great.
Apply Gartner's fivefold move. Cost per workflow goes from $0.08 to $0.40. Now 200 workflows cost $80 against $50 of revenue, and you are paying customers to use your product. Same seat price, same usage, negative unit economics.
And that is only the model invoice. The invisible stack is where the real leak lives. A Gettia Consulting breakdown of a typical production agent in 2026 lists roughly $5,500 per year in LLM inference, $2,100 in monitoring with Langfuse, $2,000 in application infrastructure, $440 for a managed vector database, and under $100 in embeddings. Add it up and you are at about $10,140, which means the model bill is barely half of what you actually spend.
Cockroach Labs published the single most useful cost insight I have seen all year on June 10, 2026: the biggest invisible cost in agentic systems is re-sent context. Every loop resends the conversation, the tool outputs, and the scratchpad. They cited Stanford Digital Economy Lab research putting re-sent context at 62% of total agent inference bills. Treat vendor-cited percentages carefully, but the direction is not in dispute.
So where is the money? Three places, and all three are stupid easy to attack.
First, prompt caching. If most of your token spend is context you already paid for once, caching is the highest-return move available and it requires no model change. Second, retry discipline. Splunk's point is that failed and repeated attempts are billed at full price, and McKinsey estimates roughly 60% of agentic cost comes from refinement, repair, and re-verification rather than the first answer. Third, model routing. Sommer's warning is blunt: there is no economical one-size-fits-all model coming, so competitive products will run multimodel ecosystems.
On pricing, the flat seat is the trap. Usage-based, credit-based, or outcome-based pricing lets your revenue move with your cost. My read on this is that flat-fee agent pricing in 2026 is the modern version of an unlimited buffet with no door count.
Three signals inside the same shift
Prices fall on a curve, complexity rises on a step function.
Splunk reported on June 29, 2026 that agentic systems burn 5 to 30 times more tokens per task than a single chatbot call, and McKinsey puts some agentic tasks near 1,000 times a simple chat exchange. Per-token prices dropped roughly 99.7% since GPT-3-era rates, but the step function wins. Your margin is the paper between the two blades.
The model invoice is barely half of what you actually spend.
A Gettia Consulting breakdown of a production agent in 2026 lists roughly $5,500 per year in LLM inference, $2,100 in monitoring, $2,000 in application infrastructure, $440 for a managed vector database, and under $100 in embeddings. Total is about $10,140. Teams that optimize only the token line are optimizing the smallest number on the bill.
Adoption is climbing straight into a cancellation wave.
Gartner survey data cited by Digital Applied shows enterprises with at least one agent in production rising from 9% in 2024 to 19% in 2025 to 31% in 2026, and Ringly.io puts 51% in production this year. Gartner also predicts more than 40% of agentic AI projects will be cancelled by the end of 2027. Instrumented unit economics is the difference between the two groups.
2031
Zoom out and the picture gets more interesting than a cost complaint.
Token price is your vendor's problem. Cost per outcome is your business. Companies that confuse the two will spend the next two years optimizing the smallest line on the invoice.
Adoption is not slowing down to wait for the math to work. Gartner survey data cited by Digital Applied shows enterprises with at least one agent in production going from 9% in 2024 to 19% in 2025 to 31% in 2026. Ringly.io puts 51% of enterprises with agents in production in 2026, with another 23% actively scaling. The buying pressure is real, which means the cost pressure arrives whether you modeled it or not.
Gartner has also predicted that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing rising costs, unclear value, and weak risk controls. Read that as a market-clearing event, not an obituary. Projects with instrumented unit economics survive. Projects running on vibes and a flat subscription do not.
The asymmetric bet here is unglamorous. Spend three weeks building a cost ledger per workflow, and you buy the option to raise autonomy later without guessing at margin. Skip it, and every capability upgrade becomes a coin flip on profitability. Amateurs ask what the model costs. Operators ask what a solved customer problem costs.
It is unclear whether the fivefold curve holds. McKinsey frames this as a management problem, arguing that better task selection and fewer refinement loops can bend the curve. Deloitte says the real constraint is foundation work: legacy integration, data architecture, and governance. Both views imply the 5x is a default outcome for teams who do nothing, not a physical law.
Only cash is real. The rest is a projection with a footnote.
What to Build This Weekend
Pick one agent workflow you already ship. Not the roadmap. One live workflow.
First, instrument it. Log input tokens, output tokens, tool calls, retries, and wall-clock time for every single run, tagged by customer. Cost per workflow is just total spend divided by successful completions, and you need the denominator to say "successful," not "attempted." Most teams cannot produce this number today, which is exactly why it is valuable.
Second, run the fivefold stress test. Multiply your current cost per workflow by five and hold your price constant. If margin goes negative, you have a pricing problem, not an engineering problem, and you have until 2028 to fix it.
Third, turn on prompt caching and cap your retry loop at three attempts with a cheaper fallback model. Then measure again. Expect things to break, because agents fail in ways unit tests do not catch. That is normal, and it is cheaper to discover on Saturday than in a customer's production account.
Some tools from today's digest that fit this loop. Grain records, transcribes, and summarizes calls with your own custom AI prompts, which is a clean way to see how much context you are actually resending in a summarization pipeline. Maskara AI runs a live debate between top models and returns the winning answer, useful when you want a router sanity check without pasting the same prompt three times. Jason AI automates B2B outreach sequences and inbound replies, and it is the perfect candidate for a one-segment test where you can measure cost per booked meeting instead of cost per token. RemNote turns notes into spaced-repetition flashcards, which is how you actually retain your own cost model instead of rediscovering it next quarter.
You do not need a CS degree to do any of this. You need a spreadsheet, one metric, and the willingness to look at a number you have been avoiding.
Turn one live agent workflow into a cost ledger you can defend.
- Instrument one shipped workflow, not the roadmap. Log input tokens, output tokens, tool calls, retries, and wall-clock time for every run, tagged by customer. Then compute total spend divided by successful completions, never attempted ones.
- Run the fivefold stress test. Multiply your current cost per workflow by 5 and hold your price constant, the way Gartner's 2028 forecast implies. If a $0.08 workflow becomes $0.40 and 200 runs per seat cost $80 against $50 of revenue, you have a pricing problem, not an engineering problem.
- Attack caching, retries, and routing in that order. Turn on prompt caching first because re-sent context is an estimated 62% of the bill, cap retry loops at three attempts with a cheaper fallback model, then measure again. Expect breakage, and be glad you found it on a Saturday instead of in a customer account.
Token price is your vendor's problem. Cost per outcome is your business.
The fivefold figure is a forecast, not an audit, and both McKinsey and Deloitte argue the curve can be bent by better task selection, fewer refinement loops, and real data and governance foundations. That is exactly the point: 5x is the default outcome for teams who do nothing, not a law of physics. Companies that ship a cost ledger per workflow buy the option to raise autonomy later without gambling on margin. Companies that price a flat seat at $50 and hope will find out from finance instead. Only cash is real, and the rest is a projection with a footnote.