Strategy

Intelligence Is Deflating 10x a Year. Your AI Contract Doesn't Know Yet.

B

Benjamin Hopwood

Operations Scaling | Agentic AI Orchestration

September 9, 2026|5 min read
Intelligence Is Deflating 10x a Year. Your AI Contract Doesn't Know Yet.

The Number Your Vendor Won't Put on a Slide

The price of AI capability is collapsing on a curve most buyers have never seen written down. Andreessen Horowitz's LLMflation tracker measured it at roughly 10x per year: output matching the original GPT-3 fell from $60 per million tokens in late 2021 to six cents by late 2024, a thousandfold drop in three years. Epoch AI found the same shape at the high end: matching the original GPT-4 on PhD-level science questions cost $37.50 per million tokens in March 2023 and $0.12 by December 2024.

If your company buys AI in any form, seats, add-ons, or committed spend, that curve is your planning assumption. Last week's news is the reason to believe it is not slowing down.

A First Chip With No Right to Be This Good

In late August OpenAI published the first results for Jalapeño, the inference chip it designed with Broadcom, and presented it at the Hot Chips conference. It went from initial design to manufacturing tape-out in nine months; custom silicon of this class normally takes two years or more. On SemiAnalysis's public InferenceX benchmark it delivered 1.5 to 1.9 times more inference work per watt than Nvidia's Blackwell-generation systems, at 700 watts against their 1,200 to 1,400. The caveats, stated plainly: it only does inference, OpenAI says it will keep buying Nvidia for training, and the comparison is against Nvidia's current generation rather than the upcoming Rubin parts, which SemiAnalysis argues would be the fairer match.

The speed came from an unusual workforce. OpenAI says its own models helped design the chip, and it showed AI-designed circuit blocks beating human baselines: a floating-point multiplier 56 percent faster, a matrix unit 10 percent smaller. A scaled-up internal version of Codex wrote the low-level kernels, including working kernels for a model architecture OpenAI had never supported, with no kernel engineer intervening; some ran 1.5 to 1.8 times faster than the human-expert versions.

One honest caveat belongs next to that. Kernel writing is the friendliest possible territory for coding agents, because a kernel's correctness can be checked by a numerical test and its quality by a stopwatch. Work with cheap, automatic verification falls to agents first. Fields where checking is expensive will fall more slowly.

Why This Reaches Nvidia's Moat Anyway

Nvidia's premium has two layers: the best hardware, and CUDA, twenty years of software with roughly six million developers on it. The lock-in was the cost of rewriting, and that labor is what agents now absorb. Jeremy Nixon, a former Google Brain engineer, says his team rebuilt CUDA-like software for the chip startup D-Matrix in about ten hours using coding agents. Credible people push back: Chris Lattner of Modular calls the moat-is-dead narrative very overblown, and Bing Xu, whose chip-software company Nvidia acquired, says verifying AI-written code is the real bottleneck. It is also true that most buyers never touch CUDA directly; they run inference through hosted APIs and serving frameworks, and Nvidia's advantages in networking, drivers, and ecosystem maturity sit above the kernel layer. Kernels are one layer of several. But the labor cost of attacking the other layers is falling on the same curve.

The Spice Trade Parallel

For about two centuries the Dutch East India Company kept nutmeg among the most expensive substances in Europe because it grew only on islands the company controlled. The premium was enforced scarcity, and when seedlings finally took root in other soil, the price collapsed within a generation. AI hardware has run on two scarcities of its own: the few thousand people who can design chips at this level, and the one software stack that ran models well. Both are leaking. Reporting in early 2025 put OpenAI's chip team at around 40 people, and the models did much of the rest. The same week as the Jalapeño results, the Chinese lab Z.ai revealed that an anonymous model which had quietly topped usage charts on Western routing platforms was its own open-weights release, served, Z.ai says, entirely on Chinese-made chips. CNBC could not verify the hardware claim. The direction does not depend on it: the number of entities that can credibly design AI silicon is about to grow from a handful to dozens, in China and elsewhere.

What Did Not Break

Nvidia is still the most valuable company on earth at about $5.3 trillion, and it reported $96 billion in quarterly revenue the day after the Jalapeño results, up 106 percent year over year. Asked about the chip on CNBC's Mad Money, Jensen Huang shrugged: lots of projects get started, lots get canceled. Training remains Nvidia's. And manufacturing stayed scarce: TrendForce reports TSMC's advanced packaging still runs 10 to 20 percent short of demand this year, a Morgan Stanley-sourced estimate puts Nvidia at roughly 60 percent of that capacity, and Samsung and SK Hynix say high-bandwidth memory is effectively sold out into 2027. Jalapeño queues for the same memory and packaging as everyone else. Designing chips got dramatically easier. Building them did not.

The Contract Playbook

Here is the part that becomes your story: wholesale AI prices fall 10x a year, but retail prices only fall when buyers push. Google folded Gemini into Workspace in early 2025, cutting a typical combined bill from $32 to $14 per user per month. OpenAI trimmed ChatGPT Business seats by about five dollars this spring. Microsoft's enterprise Copilot add-on has sat at exactly $30 per user per month since November 2023. The deflation reaches buyers who ask for it.

What to do about it, starting with the habit and then the paper:

  1. Break the "best by default" habit first. The biggest line in most AI budgets is not clause language; it is running the flagship model on work a near-frontier tier now does at a twentieth of the price, when a little workflow engineering gets better results for a fraction of the cost. Audit that before you touch the contract; the audit is its own article, coming in this series.
  2. Cap the term at twelve months. In a market deflating 10x a year, a multi-year committed spend is a bet against your own interests. OpenAI's standard terms make minimum commitments non-cancellable; keep them short.
  3. Kill the auto-renewal. OpenAI's standard order form silently renews at the same size and duration unless you give thirty days' written notice. The California State University system negotiated that line to "this Order Form will not renew," making renewal opt-in. At minimum, calendar the notice window.
  4. Get price protection in writing. CSU's negotiated clause holds its per-license fee and discounts even if it cuts seat count by up to half at renewal. That is what "negotiate like prices are falling" looks like as contract language.
  5. Get your data out before it is deleted. Standard terms promise deletion after termination; CSU added the right to export everything first. Ask for export rights, and build so that switching providers is a configuration change rather than a rebuild.

The chip news is a market-structure story. Your renewal date is where it becomes a line on your P&L.


Agentic Solutions helps companies turn shifts like this one into vendor terms, architecture choices, and budgets before the market reprices them. Start the conversation.