Future of Commerce

Backgrounder: Falling AI Inference Costs Are Rewriting Software Economics

Cheaper model output expands the field of viable AI products, but durable advantage is moving away from raw access and toward workflow, distribution, and proprietary context.

Editorial illustration for Backgrounder: Falling AI Inference Costs Are Rewriting Software Economics
Shift Signal Editorial Desk

01 · The problem

What changed

The price of using capable AI models has fallen fast enough to change which products can exist. Stanford's 2025 AI Index estimated that the inference cost for a system reaching roughly GPT-3.5-level performance on one benchmark fell from about $20 per million tokens in November 2022 to $0.07 by October 2024—a decline of more than 280 times [2].

That is not merely a cheaper technology bill. It changes the minimum revenue a product needs before an AI-assisted workflow can make economic sense. Features that once required expensive, carefully rationed calls can move into routine search, drafting, support, coding, document review, and operations.

02 · The stakes

Why it matters

The trap is assuming that falling input costs automatically create durable margins. They also lower the barrier for competitors. If two teams can buy similar model capability from the same handful of providers, access to the model is rarely the defensible part of the business.

Cheap inference can also encourage waste: excessive prompts, duplicated context, weak evaluation, and features that look impressive but do not improve an outcome customers will pay for. A lower unit price does not rescue a workflow with poor retention or no measurable return.

03 · The evidence

What the record shows

The technical break began before the commercial cost curve. The 2017 “Attention Is All You Need” paper described a model architecture built around attention rather than recurrence or convolution [1]. That design made training more parallelizable and became a foundation for the systems now sold through commercial APIs.

The adoption record shows both acceleration and limits. Stanford's 2026 AI Index reports broad organizational adoption of AI and generative AI, while deployment of agents remains in the single digits in most surveyed business functions [3]. In other words, buying model access is common; redesigning a dependable end-to-end process is not.

The useful business distinction is therefore between a demonstration and an operating system. A demonstration produces a striking output once. An operating system has inputs, permissions, evaluation, human escalation, cost controls, and a feedback loop that improves the work.

04 · The response

What to do

Operators should measure AI features at the task level. Track model cost per completed outcome, error and escalation rates, time saved, and the percentage of outputs actually used. Then compare those numbers with the non-AI process.

Products should also be designed so model vendors can change. Keep business rules, customer context, evaluation cases, and workflow state outside any single provider's prompt format. That architecture does not eliminate switching costs, but it prevents the model endpoint from becoming the product's only source of value.

05 · The bigger signal

What to watch next

As inference becomes cheaper, the scarce assets move up the stack: trusted customer relationships, permissioned data, domain-specific evaluation, reliable integrations, and distribution. The winners may spend less time advertising that they “use AI” and more time proving that a particular job is completed faster, more safely, or at lower total cost.

Watch the gap between adoption and autonomy. Broad experimentation can grow quickly while dependable agent deployment remains hard. Companies that close that gap with rigorous workflow design—not merely more model calls—have the better chance of keeping the value created by the falling cost curve.

Action desk

Your next moves

  1. 01

    Choose one repeated workflow and record its current time, cost, and error rate before adding AI.

    Time: 30 minutes

  2. 02

    Create a small evaluation set from real, permissioned work and rerun it whenever the model or prompt changes.

    Time: 2 hours

  3. 03

    Separate proprietary context and workflow state from the model provider so the application can switch models deliberately.

    Time: 1 day

Evidence

Sources

3 cited

  1. [1]
    Attention Is All You Need

    arXiv · Primary source

  2. [2]
    AI Index Report 2025

    Stanford Institute for Human-Centered Artificial Intelligence · Primary source

  3. [3]
    2026 AI Index Report: Economy

    Stanford Institute for Human-Centered Artificial Intelligence · Primary source

Disclosure: Backgrounder published on its actual publication date using previously released primary sources. It is not presented as contemporaneous coverage of earlier events.

Pass the signal

Know someone this affects?

Download graphicLinkedInXFacebook

Get the next briefing

One useful signal, delivered daily.

Keep up with platform changes, earnings pressure, and practical next moves.