Future of Commerce
Backgrounder: Falling AI Inference Costs Are Rewriting Software Economics
Cheaper model output expands the field of viable AI products, but durable advantage is moving away from raw access and toward workflow, distribution, and proprietary context.

01 · The problem
What changed
The price of using capable AI models has fallen fast enough to change which products can exist. Stanford's 2025 AI Index estimated that the inference cost for a system reaching roughly GPT-3.5-level performance on one benchmark fell from about $20 per million tokens in November 2022 to $0.07 by October 2024—a decline of more than 280 times [2].
That is not merely a cheaper technology bill. It changes the minimum revenue a product needs before an AI-assisted workflow can make economic sense. Features that once required expensive, carefully rationed calls can move into routine search, drafting, support, coding, document review, and operations.
02 · The stakes
Why it matters
The trap is assuming that falling input costs automatically create durable margins. They also lower the barrier for competitors. If two teams can buy similar model capability from the same handful of providers, access to the model is rarely the defensible part of the business.
Cheap inference can also encourage waste: excessive prompts, duplicated context, weak evaluation, and features that look impressive but do not improve an outcome customers will pay for. A lower unit price does not rescue a workflow with poor retention or no measurable return.
03 · The evidence
What the record shows
The technical break began before the commercial cost curve. The 2017 “Attention Is All You Need” paper described a model architecture built around attention rather than recurrence or convolution [1]. That design made training more parallelizable and became a foundation for the systems now sold through commercial APIs.
The adoption record shows both acceleration and limits. Stanford's 2026 AI Index reports broad organizational adoption of AI and generative AI, while deployment of agents remains in the single digits in most surveyed business functions [3]. In other words, buying model access is common; redesigning a dependable end-to-end process is not.
The useful business distinction is therefore between a demonstration and an operating system. A demonstration produces a striking output once. An operating system has inputs, permissions, evaluation, human escalation, cost controls, and a feedback loop that improves the work.
04 · The response
What to do
Operators should measure AI features at the task level. Track model cost per completed outcome, error and escalation rates, time saved, and the percentage of outputs actually used. Then compare those numbers with the non-AI process.
Products should also be designed so model vendors can change. Keep business rules, customer context, evaluation cases, and workflow state outside any single provider's prompt format. That architecture does not eliminate switching costs, but it prevents the model endpoint from becoming the product's only source of value.
05 · The bigger signal
What to watch next
As inference becomes cheaper, the scarce assets move up the stack: trusted customer relationships, permissioned data, domain-specific evaluation, reliable integrations, and distribution. The winners may spend less time advertising that they “use AI” and more time proving that a particular job is completed faster, more safely, or at lower total cost.
Watch the gap between adoption and autonomy. Broad experimentation can grow quickly while dependable agent deployment remains hard. Companies that close that gap with rigorous workflow design—not merely more model calls—have the better chance of keeping the value created by the falling cost curve.
Action desk
Your next moves
- 01
Choose one repeated workflow and record its current time, cost, and error rate before adding AI.
Time: 30 minutes
- 02
Create a small evaluation set from real, permissioned work and rerun it whenever the model or prompt changes.
Time: 2 hours
- 03
Separate proprietary context and workflow state from the model provider so the application can switch models deliberately.
Time: 1 day
Evidence
Sources
3 cited
- [1]Attention Is All You Need
arXiv · Primary source
- [2]AI Index Report 2025
Stanford Institute for Human-Centered Artificial Intelligence · Primary source
- [3]2026 AI Index Report: Economy
Stanford Institute for Human-Centered Artificial Intelligence · Primary source
Pass the signal
Know someone this affects?
Get the next briefing
One useful signal, delivered daily.
Keep up with platform changes, earnings pressure, and practical next moves.