The news
AI model pricing is no longer moving in one direction. In September 2026, OpenAI, Anthropic and Alibaba (BABA) cut prices and a Japanese start-up launched a cheap router, while DeepSeek, long the low-price benchmark, is reportedly growing revenue on rates it raised in August.
The biggest cuts to text-model prices came on September 22. OpenAI's developer changelog lists its new GPT-6 Sol reasoning model at $2 per million input tokens and $10 per million output tokens, and GPT-6 Luna at $0.10 and $0.50. Tech Wire Asia reported those rates are half or less of the prior GPT-5.6 Sol and Luna prices. A token is a small chunk of text and the unit AI providers bill by.
Anthropic's Claude Opus 5.5, launched the same day, costs $4 per million input tokens and $20 per million output, down from $5 and $25 for Opus 5, and Anthropic says it costs 40% less on typical workloads. On September 28, Anthropic kept Sonnet 5.5 at the same $2 and $10 as Sonnet 5 but said the model needs fewer tokens, cutting cost per task by up to 30%.
Cheaper options arrived from elsewhere too. On September 11, Sakana AI released Fugu Max, a service that routes each request across many models, including open-weight ones, at $2 per million input tokens and $6 per million output. Alibaba's Qwen-Audio 3.1 release on September 23 cut speech-recognition prices by up to 95%, text-to-speech by about 70% and real-time voice by about 85%, The Decoder reported.
DeepSeek went the other way. VentureBeat reported in August that the Chinese developer raised prices steeply, with increases for its Flash and Pro models of up to 371% and 355% respectively as announced at the time; DeepSeek's current Flash rates are slightly lower. Its pricing page now charges twice the off-peak rate during weekday peak hours, 01:00 to 04:00 and 06:00 to 10:00 UTC. National Technology News, citing The Information, reported on September 24 that DeepSeek's annualized revenue has reached $1 billion, roughly double earlier in the year, as it seeks about $7.5 billion in new funding.
The numbers
- GPT-6 Sol list price (input / output per million tokens)
- $2 / $10
- Claude Opus 5.5 list price (input / output per million tokens)
- $4 / $20, down from $5 / $25
- Sakana Fugu Max price (input / output per million tokens)
- $2 / $6
- Alibaba speech-recognition price cut, per The Decoder
- Up to 95%
- DeepSeek peak-hour premium over off-peak
- 2x
- DeepSeek annualized revenue, per The Information via National Technology News
- $1 billion
Why CEOs should care
CFOs should stop budgeting AI on per-token list prices alone. Anthropic says Sonnet 5.5, at unchanged rates, can cost less per job because it uses fewer tokens, while DeepSeek shows the cheapest supplier can reprice quickly once demand is proven. The number that matters is cost per completed task, measured on your own workloads. Ask finance and engineering to agree on that metric and re-run it whenever a vendor changes models or rates.
Procurement teams should write price movement into contracts. Seek clauses that pass list-price cuts through to committed spend, cap increases over the term, and allow a switch to successor models at no higher cost. For time-sensitive workloads, check whether a provider uses time-of-day pricing, as DeepSeek now does, and whether batch jobs can shift to cheaper hours. Routers such as Sakana's Fugu Max also mean procurement may be buying access to many models at once, which needs a clear list of approved underlying providers.
For CIOs and CISOs, the push to chase the lowest price raises governance questions. Moving workloads between US, Chinese and multi-model routing services changes where data goes and under which laws. Keep a register of which providers touch which data, and require the same security review for a router as for any direct model supplier. Boards should see the upside too: OpenAI's chief financial officer, Sarah Friar, said an 80% price cut for its Luna model in July helped drive roughly a tenfold rise in usage, Tech Wire Asia reported, so lower prices can unlock projects that were uneconomic months ago.
The bigger picture
The split reflects different business positions. US labs are using price cuts on new models to win volume, while DeepSeek, if reports of its revenue growth are accurate, is showing that a low-cost provider with strong demand can raise rates and grow revenue at the same time. The net result for buyers is more choice and more volatility, and advantage shifts to organizations whose software can move a workload from one model to another without a rebuild.
What’s next
Expect more repricing as other providers respond to the September 22 cuts and as new models arrive. Watch whether DeepSeek's reported fundraising closes, whether other providers adopt peak-hour pricing, and whether OpenAI and Anthropic extend cuts to older models that many enterprises still run in production.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error









