The news
Cribl and CData each unveiled an AI gateway on September 29, software that sits between a company's apps and AI models to route prompts, enforce budgets and control data access, as agents make per-token AI bills harder to predict.
San Francisco-based Cribl announced StreamAI, an enterprise AI gateway with a built-in model router that sends each prompt to a model based on performance, cost and context. The company said built-in circuit breakers let IT teams limit token use, and when an app hits a cost limit, StreamAI switches to an alternative model instead of stopping. It also redacts sensitive data and logs each routing decision for audits.
Cribl said it is offering free inference to customers who use automatic routing to AI models it has benchmarked. It did not state limits, pricing or a launch date, saying only that StreamAI will be available soon and that customers can ask their account team for early access. In Cribl's own SecIT Bench test of 20 models on 30 real IT and security investigations, the company found a 17% spread in diagnostic accuracy but a 20x range in what each investigation cost. "The right model depends on the job, the context, and the economics," said co-founder and CEO Clint Sharp.
CData, based in Chapel Hill, North Carolina, opened an early access program for Connect AI Gateway. It gives IT teams one place to register AI models, agents and MCP servers (a common way for agents to plug into business tools and data), then set rules for what each can reach, down to individual records. CData said spending is attributed and budgeted by team, agent, model and tool, and each request goes to the most efficient model that can handle it. Pricing was not disclosed.
CData also said that in a study of 22 models run against live business systems, every model returned the same correct answer through its governed tools, and the cheapest did so at 178 times lower cost than the most expensive. SiliconANGLE, which reported the gap as up to 175-fold, noted that these are company benchmarks, not independent tests or savings documented at a customer. CData chief marketing officer Will Davis told the outlet that routing will be policy-based at first, with automatic selection of the lowest-cost model to come later.
The numbers
- Cost gap, cheapest vs priciest model with same correct answer (CData study, 22 models)
- 178x
- Spread in diagnostic accuracy across 20 models (Cribl SecIT Bench)
- 17%
- Range in investigation spend across the same models (Cribl)
- 20x
- Token use per task, agentic AI vs a simple call (Futurum)
- 10 to 100 times
- Forecast total inference spending by 2030 (Futurum)
- $885 billion
Why CEOs should care
For CFOs, AI is moving from a predictable per-seat license to a meter. Futurum research cited by ZK Research analyst Zeus Kerravala in SiliconANGLE says agentic AI can drive token use per task 10 to 100 times higher than a simple AI call. Before a gateway is bought, decide who owns it: finance, IT, security or the AI team. Then ask any vendor whether spend can be tracked by team and by agent, what a hard cap looks like, and what happens when an app hits it.
That last question matters for quality as well as cost. Cribl's design falls back to another model when a budget runs out, so buyers should test whether answers from the cheaper fallback are good enough for the work in question. Cribl's own benchmark found accuracy varied far less than cost, but that was measured on IT and security investigations, not on every task a company runs.
For CISOs, a gateway is also a checkpoint for data. Cribl says StreamAI redacts sensitive data; CData says it enforces permissions down to the record level. Ask where prompts and logs are stored, whether the audit trail shows which prompt, model and tool touched which record, and what happens to AI traffic if the gateway itself goes down. Every AI request would now pass through one more vendor.
The bigger picture
The spending pressure behind these launches is large. Kerravala cites Futurum forecasts that total inference spending will rise from $120 billion in 2025 to $885 billion by 2030, and that agent and reasoning inference will grow 219% in 2026. A Futurum report sponsored by bare-metal cloud provider QumulusAI found reserved and owned infrastructure make up 66% of AI compute use, compared with 19% for on-demand cloud. Kerravala argues per-token pricing suits experiments but gets costly once agents run in production.
Gateways offer a middle path: keep buying tokens, but route each request to the cheapest model that can do the job and stop spending at a set limit. Both vendors' savings figures come from their own tests, so companies should run a trial on their own workloads before treating those numbers as a forecast.
What’s next
CData's early access program began on September 29, and Cribl is taking early access requests through its account teams. Watch for published pricing, the terms and limits of Cribl's free inference offer, and whether independent tests confirm the cost gaps both companies report.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error







