Dev.to AI 🤖 Ai 👁 0 📖 5 min read

AI Inference Costs Are Falling 13x a Year: What the Price Collapse Means for Builders

On January 31, 2025, OpenAI's o3 cost about 30 cents per question to score 75% on GPQA Diamond, a PhD-level science benchmark. Just under 18 months later, GPT-5.6 Luna matched that score for $0.0004 - a 725-fold drop in

AI Inference Costs Are Falling 13x a Year: What the Price Collapse Means for Builders

On January 31, 2025, OpenAI's o3 cost about 30 cents per question to score 75% on GPQA Diamond, a PhD-level science benchmark. Just under 18 months later, GPT-5.6 Luna matched that score for $0.0004 - a 725-fold drop in the price of thought. (Source: Epoch AI, 2026)

Infographic

The Collapse, Measured

Epoch AI tracked five benchmarks across three years of model releases: for a fixed level of capability, the cheapest price fell about 47% per quarter since 2023, or 13 times per year. (Source: Epoch AI, 2026)

That pace leaves every comparable technology behind: four times faster than DNA sequencing, six times faster than computing, 18 times faster than lithium batteries, and 54 times faster than electricity in the century through 1973. (Source: Epoch AI, 2026)

The decline depends on the task. Across six benchmarks, annual decline ranged from 9x to 900x, with GPT-4-level performance on PhD science questions getting 40x cheaper per year. (Source: Epoch AI, 2025)

The discount also has a lifespan. On average, the cost to reach a capability level falls 66% per quarter while that level is state of the art, then slows to 32% per quarter two years later. (Source: Epoch AI, 2026)

Cheap Tokens Do Not Mean Cheap Agents

Gartner projects that by 2030, inference on a trillion-parameter model will cost over 90% less than in 2025, and models will be up to 100 times more cost-efficient than their 2022 equivalents. (Source: Gartner, 2026)

Agents break the simple arithmetic. Agentic models consume 5 to 30 times more tokens per task than a standard chatbot, and they run vastly more tasks. Gartner expects total inference costs per agentic workflow to increase more than fivefold through 2028, even as per-token prices keep falling. (Source: Gartner, 2026)

Will Sommer of Gartner put it plainly: product leaders "should not confuse the deflation of commodity tokens with the democratization of frontier reasoning." (Source: Gartner, 2026)

The Hardware Counter-Movement

A parallel push is moving the economics onto hardware people already own. Magnitude, an open-source inference engine, compiles and tunes its kernels on the target device before a model runs, measuring up to 2x faster than llama.cpp, with 92% faster decoding on Apple Silicon and 19% on NVIDIA GPUs. (Source: Magnitude, 2026)

Model compression is closing the gap from the other side. The Strata project runs Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model, on a single 12-24 GB NVIDIA card plus 64 GB of RAM, writing at 60 to 95 tokens per second; on an RTX 5070 it measured 93 tokens per second. (Source: Strata, 2026)

A capability that wanted a server rack in 2024 now fits under a desk, a shift that rewrites which experiments smaller teams can afford.

What This Looks Like in the Philippines

Philippine demand is large and lopsided. The US International Trade Administration values the country's AI market at roughly $772 million in 2024, on pace for $3.5 billion by 2030, a 28.6% compound annual growth rate. (Source: US ITA, 2026)

Adoption splits the same way: about 67% of IT-BPM firms have deployed AI tools, against 14.9% of Philippine firms overall. (Source: US ITA, 2026; PIDS, 2025)

Consumers run ahead of enterprises. The Philippines ranks sixth worldwide for ChatGPT usage, with 42.4% of internet users having used it in the past month against a 26.5% global average. (Source: Digital in Asia, 2026)

Policy is responding: the Department of Trade and Industry's National AI Strategy Roadmap 2.0, released in July 2024, prioritizes AI in healthcare, education, agriculture, logistics, and digital services. (Source: US ITA, 2026)

For the 99.6% of Philippine establishments that are MSMEs, the falling price of inference changes which workflows are worth automating first, not whether. (Source: DTI, 2025)

The Adoption Gap Prices Do Not Fix

Cheap inference expands what is technically possible. It does not manufacture trust. A February 2026 Pew survey found 51% of Americans avoided AI chatbots entirely, with 79% of that group citing privacy concerns. (Source: Axios, 2026)

The resistance is sharpest where agents need real access. In a Thales digital trust survey, only 13% of respondents would let an AI helper read their email, 11% would let one rebook travel, and 7% would let one move money between bank accounts. (Source: Axios, 2026)

Where adoption exists, it is deep but narrow. The Federal Reserve Bank of St. Louis describes it as widespread but shallow: at least 20% of workers use AI in more than 80% of occupations, heavily weighted toward white-collar and tech-adjacent roles. (Source: Axios, 2026)

FAQ

Q: How fast are AI inference costs actually falling?
A: For a fixed level of capability, about 47% per quarter since 2023, or roughly 13x per year, according to Epoch AI's analysis of five benchmarks. (Source: Epoch AI, 2026)

Q: If tokens keep getting cheaper, why are AI budgets rising?
A: Volume. Agentic workflows consume 5 to 30 times more tokens per task than chatbots, and Gartner projects total inference costs per agentic workflow will grow more than fivefold through 2028. (Source: Gartner, 2026)

Q: Can large open models run on ordinary hardware now?
A: Yes, with limits. Strata runs a 125-billion-parameter open-weight model on a 12-24 GB consumer GPU plus 64 GB of RAM at 60-95 tokens per second, and Magnitude tunes models to run up to 2x faster than llama.cpp on hardware people already own. (Source: Strata, 2026; Magnitude, 2026)

Key Takeaway

The price of a unit of AI capability is collapsing faster than any general-purpose technology on record, and that changes what is worth building, not just what is cheap to run. The trap is reading cheap tokens as a strategy: Gartner's own forecast says frontier reasoning stays scarce and agentic bills rise through 2028. The builders who win the next year will route routine work to small models, gate expensive reasoning, and design for trust rather than for token budgets. Yano.AI builds multi-agent systems and reads the same data the same way, because the constraint that survives every price drop is the process the model has to connect to.

So here is the question worth answering this week: if a unit of intelligence costs 725 times less than it did 18 months ago, which task in your business are you still pricing at last year's rates?

Sources

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.