The news
DeepInfra, a Palo Alto company that runs open-source AI models for other businesses, said on September 29, 2026 that its annual recurring revenue (ARR) has passed $100 million. It put its annualized run-rate at more than $105 million, nearly 14 times a year earlier.
The company also said it now processes 22 trillion tokens a week and has more than tripled its token volume since May. Tokens are the small chunks of text that AI models read and write, and their volume is a rough measure of how much work a model provider is doing for customers.
DeepInfra sells inference, the step where a trained model is actually run to answer requests, rather than training models of its own. The company says it supports more than 200 open-source models through application programming interfaces (APIs) compatible with OpenAI's. In May, SiliconANGLE reported the figure as more than 190 models.
The announcement named three customers: LiveKit, Humans& and OpenCode. LiveKit's inference lead, Neil Dwyer, said in the release that LiveKit needs time-to-first-token below 500 milliseconds and output above 100 tokens per second with no downtime, and that DeepInfra has helped it scale to many billions of tokens per day.
The milestone comes months after DeepInfra's $107 million Series B. SiliconANGLE reported on May 4, 2026 that 500 Global and Georges Harik, whom it described as one of Google's first cloud engineers, co-led that round, with participation from Nvidia (NVDA), Samsung Next, Supermicro (SMCI), Felicis and others. No valuation was disclosed then or in the new announcement.
Co-founder and chief executive Nikola Borisov said in the release that AI spending is moving "decisively from the lab into production." The company also said it recently appointed Behzad Nouri to lead go-to-market, and that it has added its first data center space in Canada.
The numbers
- Annualized run-rate
- More than $105 million
- Year-over-year growth (company figure)
- Nearly 14x
- Weekly token volume
- 22 trillion tokens
- Series B (May 2026)
- $107 million
- Open-source models supported
- 200+
Why CEOs should care
For technology buyers, the growth of a provider like DeepInfra is a sign that running open-source models through a specialist is now a mainstream option, not a science project. If your teams are building on closed-model APIs or a single large cloud, the practical move is to benchmark: run the same workload on an open model through two or three inference providers and compare cost per million tokens, latency and uptime. LiveKit's stated targets, under 500 milliseconds to the first token and more than 100 tokens per second, are a useful template for writing your own service levels into a contract.
CFOs should read the headline figure carefully. An annualized run-rate takes a recent period of revenue and projects it over a year; it is not a year of booked revenue, and usage-based inference revenue can swing as customers shift workloads. Growth of nearly 14 times in a year also means most of that revenue is recent. Ask any provider for volume commitments, price protection and notice periods, and model what happens if you need to move workloads quickly.
CISOs and compliance teams have concrete items to check. DeepInfra says it offers zero data retention and holds SOC 2 and ISO 27001 certifications, and it now has data center space in Canada as well as the United States. Buyers should confirm in writing where prompts are processed, whether logs are kept, and which subprocessors touch the data, especially for regulated or customer data.
The bigger picture
The announcement fits a broader shift of AI spending toward running models in production, the point Borisov makes. SiliconANGLE reported in May that DeepInfra operates its own hardware across eight U.S. data centers instead of renting spare capacity, uses Nvidia's Blackwell and Vera Rubin chips, and said more than 30% of its token volume then came from autonomous agents. The company claimed at the time up to 20 times better inference cost efficiency, without naming a baseline. Nvidia's stake in DeepInfra also shows how the chipmaker is backing the independent clouds that buy its hardware.
What’s next
Watch whether DeepInfra discloses larger named customers, new regions beyond the United States and Canada, or another funding round to pay for capacity. For buyers, the next useful data point is price: whether rising volume at independent providers leads to lower per-token rates for open-source models.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error
Companies in this story








