Listen to this newsletter ⬆️

Subscribe Forward this edition
The Agentic Enterprise
AK · Tuesday Edition · 7 min read
Tuesday, August 4, 2026
The price of frontier AI keeps falling. Your bill keeps rising.
On July 30, OpenAI cut its GPT-5.6 tiers again, with Luna now at $0.20 per million input tokens. Across vendors, the frontier token-price index sits 88% below its 2023 base.
And yet 73% of enterprises overran their 2026 AI budget. The gap between falling prices and rising bills is the real story, and it is an architecture problem, not a procurement one. The lever is no longer which model you buy. It is how you route the work.
The Big StoryMarkets / Pricing
Intelligence is commoditizing at the unit, not the invoice.
Price cuts are now the industry's steadiest drumbeat. On July 30, OpenAI trimmed its GPT-5.6 line again: Terra to $2 per million input tokens and $12 output, Luna to $0.20 and $1.20, with a new Fast mode added for the Sol flagship at $5 and $30. It is not alone. Anthropic's Claude Sonnet 5 is running an introductory $2 and $10 through August 31, Google's Gemini 3.6 Flash sits at $1.50 and $7.50, and xAI's Grok 4.5 undercuts the flagships at $2 and $6. The frontier token-price index is now 88% below its March 2023 base.

Look across the tiers and the pattern is clear. At the top, four vendors cluster within a few dollars for frontier capability. In the middle, balanced models have converged near $2 in and $10 out. At the bottom, high-volume tiers have dropped under a dollar, with open-weight options like DeepSeek billing at pennies. Capability that carried a premium a year ago is now a commodity with four suppliers.

Here is the enterprise read. Falling unit prices do not lower your bill. They raise your usage. Most enterprises overran their AI budget this year even as per-token costs dropped, because agentic workloads generate far more tokens than the spreadsheets assumed, and reasoning models bill for thinking you never see. The lever that matters is no longer the sticker price. It is routing: sending each task to the cheapest tier that can do it, and metering what your agents actually consume.

The price of a token is falling faster than almost anything in tech. The size of your bill is going the other way.
The Spearhead Take
Stop shopping for the cheapest model and start building the router. Tier your traffic: frontier models for the few tasks that need them, mid-tier for the bulk, value models and open weights for volume. The savings are not in the vendor you pick. They are in never sending a $30 output token to do a $1.20 job.
What Frontier Intelligence Costs Now
Four vendors, three tiers, one direction.
Standard API list prices, US dollars per million tokens, as of August 4, 2026. Batch APIs cut these roughly 50%; cache hits cut input up to 90%.
Frontier tier
ModelInput / 1MOutput / 1M
Anthropic Claude Fable 5$10.00$50.00
OpenAI GPT-5.6 Sol$5.00$30.00
Anthropic Claude Opus 5$5.00$25.00
Google Gemini 3.1 Pro$2.00$12.00
xAI Grok 4.5$2.00$6.00
Balanced / mid tier
ModelInput / 1MOutput / 1M
Google Gemini 3.6 Flash$1.50$7.50
OpenAI GPT-5.6 Terra$2.00$12.00
Anthropic Claude Sonnet 5 (intro, to Aug 31)$2.00$10.00
xAI Grok 4.3$1.25$2.50
Value / high-volume tier
ModelInput / 1MOutput / 1M
Anthropic Claude Haiku 4.5$1.00$5.00
Google Gemini 3.5 Flash-Lite$0.30$2.50
OpenAI GPT-5.6 Luna$0.20$1.20
xAI Grok 4.1 Fast$0.20$0.50
DeepSeek V3.2 (open weight)$0.14$0.28
The Obvious & The Overlooked
Three reads the market has made. Four it has not.
The Obvious
Model prices keep falling.
The frontier token-price index sits 88% below its 2023 base. AI Magicx
Frontier capability now has four suppliers.
OpenAI, Anthropic, Google, and xAI all cluster within a few dollars. CloudZero
Every vendor now sells a sub-dollar tier.
High-volume models have all fallen below a dollar per million input tokens. Inference.net
The Overlooked
Falling prices raised bills, not lowered them.
73% of enterprises overran their 2026 AI budget as volume outran the price cuts. The Daily Brief
Reasoning models bill for invisible thinking tokens.
The tokens you never see are the ones that blow the budget. CloudZero
The premium tier is widening, not closing.
Fable 5 at $10/$50 shows vendors still charging 5-10x for the very top. CloudZero
Open weights reset the floor.
DeepSeek V3.2 at $0.14/$0.28 makes pennies the new baseline for commodity work. Inference.net
Moving Pieces
Five developments worth a CIO's attention.
Policy / Governance
Two capitals wrote opposite rulebooks in one week

The EU AI Act's enforcement powers for general-purpose models went live August 2: the AI Office can now demand documentation, compel model access for pre-release evaluation, and fine a provider up to 15 million euros or 3% of worldwide turnover. Three days later the White House presented OpenAI, Anthropic, Google, and Meta a voluntary US framework, opt-in, capped at 30 days of early access, and explicitly barred from becoming a licensing regime. Same models, opposite postures. The model at the center of your workflow now answers to a hard regime in Brussels and a handshake in Washington, and that divergence is a procurement variable, not a legal footnote.

Sources: CNBC · Bloomberg · The Next Web
Deals / Infrastructure
Anthropic anchored a $10B compute deal to a months-old startup

Anthropic agreed to buy computing capacity in a $10 billion deal from Volta Infra Holdings, a cloud startup backed by Nvidia and only months old, Bloomberg reported August 4. The read is supply, not novelty. Compute scarcity is acute enough that a frontier lab will anchor a ten-figure contract to a company with almost no track record, as long as it has Nvidia chips and the power to run them. That is exactly the capacity your model vendor needs to serve you, and it is being locked up years in advance. It is also the cost base underneath every price cut in the story above.

Sources: Bloomberg
Deals / Security
Zenity raised $125M to police AI agents in production

Zenity closed a $125 million Series C led by Norwest, with SoftBank Vision Fund 2, Hitachi Ventures, LG Technology Ventures, and Intel Capital participating, bringing it to $185 million total. The pitch is securing autonomous agents already running inside large organizations. The timing is the story. Capital is flowing to agent security precisely because enterprises have moved past pilots into production, where a misconfigured agent is a live liability, not a demo bug. When the guardrail vendors raise nine figures, it tells you where the deployments already are.

Sources: SiliconANGLE · Fortune
Research
Gartner says quantum will not run production AI this decade

Gartner said no enterprise AI workload at scale will run on quantum hardware through 2028, and classical accelerated computing will keep winning every production benchmark for the rest of the decade. This is useful because it kills a pitch. If a vendor is selling quantum-accelerated AI as a near-term enterprise capability, Gartner just told you it is not real on your planning horizon. Spend the budget on the accelerators that actually serve models today. The quantum line item can wait until the next planning cycle, and probably the one after that.

Sources: Gartner
Workforce
AI is now the most-cited reason for layoffs

AI, automation, or machine learning was named in 54% of this year's layoff events, and cited in 101,743 US job cuts through June, nearly double the 54,836 attributed to it in all of 2025. Oracle's 30,000-person reduction leads the year. Read the framing carefully. "Because of AI" is doing double duty in these announcements: part genuine automation, part cover for ordinary cost cutting dressed in a better story. The number is large and rising either way. The fact that AI is now the preferred reason to give is itself a management signal worth tracking.

On the Radar
Nine signals, sharpened.
PricingDeepSeek V3.2 is the cheapest API in the comparison at $0.14/$0.28. The open-weight model resets the floor for commodity work and pressures every paid value tier above it. Inference.net
PricingLong prompts quietly cost double. GPT-5.6 requests above 272K input tokens bill at 2x input and 1.5x output for the whole request, a hidden-cost trap in long-context agents. DevTk.AI
PolicyThe EU AI Office is already in contact with OpenAI. OpenAI confirmed it is talking to the AI Office after the recent model cyber incidents, an early live test of Europe's new inspection powers. CNBC
DealsFireworks AI raised about $1.5B for inference. One of the year's largest rounds went to serving infrastructure, underscoring that inference, not training, is where enterprise spend now concentrates. Qubit Capital
Deals8090 Solutions raised $135M led by Salesforce. Chamath Palihapitiya's platform coordinates fleets of agents to build enterprise software, a bet on agents writing the apps. New Market Pitch
DealsHush Security raised a $30M Series A. The round, with Akamai as a strategic investor, funds identity and access control for the non-human workforce of agents and bots. AI Agent Store
DealsNatural raised $30M to be a "Stripe for AI agents." The Series A funds transaction rails so agents can pay and be paid, bringing the startup to $40M total. AI Agent Store
ProductGoogle is pulling Gemini 3.5 Flash from its global region today. The August 4 removal is one more reminder of how fast the model layer under enterprise workflows churns. Google Cloud
DeploymentBanking and insurance lead agent production at 47%. Regulated industries, not tech, are furthest into shipping agents, per 2026 adoption data. DigitalApplied
The Number
88%
Fall in the frontier token-price index since 2023
Frontier-grade capability, the most expensive thing in the stack, has commoditized as fast as the cheap tiers.
The index sat at 12 on August 3, against a base of 100 three years earlier. And yet enterprise AI bills went up this year, because usage grew faster than price fell. Cheaper intelligence is not cheaper AI. It is more AI at a lower unit price, which is a different budget entirely.
Source: BenchLM.ai
Counter-Signal
Risk / Economics
The cheap-model story is how budgets get blown.

The tidy read is that models keep getting cheaper, so AI is getting cheaper. It is not, and the confusion is expensive. Two reasons. First, a lower price per token invites far more tokens. Agentic loops, retries, long context, and reasoning traces multiply consumption faster than rates fall, which is exactly why 73% of enterprises overran budget this year while every vendor was cutting prices.

Second, the sticker rate hides the real bill. Requests over a context threshold bill at 2x, reasoning models charge for thinking tokens you never see, and cache writes cost more than cache reads. The list price is a marketing number. Your effective cost per useful task is the one that matters, and it does not drop just because a press release cut a headline rate. Meter the workload, not the rate card.

From the Field
For two years the AI cost question was simple: can we afford the frontier model? This week it flipped. The frontier model is cheap now. The real question is whether you can afford how much of it you are using.

That is a harder question, because it does not have a vendor to answer it. When the sticker price falls, the instinct is to relax and let usage grow, and that is precisely when the invoice runs away from you. The teams keeping their AI budgets flat this year are not the ones who found the cheapest model. They are the ones who built the routing layer: a cheap model handling the bulk of the work, a frontier model reserved for the few tasks that actually need it, and a meter on every agent so nobody discovers the bill at month end. The pricing page is the vendor's decision. The consumption is yours.

The vendors will keep cutting the sticker price. Whether that shows up in your budget is a choice you make in your architecture, not one they make on their pricing page.
Let's get to production,
AK
Talk to SpearheadForward this edition
The Agentic Enterprise
Know more about AI than 95% of your peers. By 7 AM.
A daily AI intelligence briefing for enterprise leaders, published by Spearhead. We build AI systems that work. Strategy. Engineering. Production. Outcomes.

Keep Reading