The news
OpenAI's GPT-6 Astra Ultrafast, a faster tier of the company's top model, runs on Nvidia (NVDA) Blackwell GPUs and generates tokens up to 8x faster than Astra's Standard mode, Nvidia said in a blog post on October 1, 2026. The tier targets coding, tool use and agent workflows.
According to Nvidia, Ultrafast is available in the OpenAI API and to eligible ChatGPT Work and Codex users. Nvidia pitched the tier at interactive applications and agentic AI, where a model may call tools and generate code many times before a person sees a result.
Two OpenAI executives were quoted in Nvidia's post. Uday Ruddarraju, described as OpenAI's CTO of compute, said OpenAI used its internal models to optimize inference on Nvidia GPUs. Philippe Tillet, OpenAI's inference lead, credited Nvidia's tooling and documentation for helping OpenAI's models program Blackwell and Rubin GPUs.
Nvidia's post did not list prices. A pricing breakdown published by OrcaRouter, an AI model routing service, citing OpenAI's API pricing page, lists Ultrafast at $60 per million input tokens, $6 for cached input, $75 for cache writes and $300 per million output tokens for requests up to 272,000 input tokens. That is six times the Standard tier, OrcaRouter said. Requests above 272,000 input tokens are repriced in full at $120 input and $450 output per million tokens.
OrcaRouter also reported that Ultrafast launched with low rate limits: 500,000 tokens per minute for usage tiers 1 to 3, 1 million for tier 4 and 5 million for tier 5. It said Ultrafast supports only US data residency and global processing, with no EU or other regional processing endpoints, and said Ultrafast moved from a waitlist to general availability on September 29, 2026.
The numbers
- Speed vs Astra Standard (Nvidia)
- Up to 8x
- Price multiplier vs Standard (OrcaRouter)
- 6x
- Input / output price, ≤272K tokens
- $60 / $300 per million tokens
- Input / output price, >272K tokens
- $120 / $450 per million tokens
- Rate limit, tiers 1-3
- 500,000 tokens per minute
Why CEOs should care
For engineering leaders building real-time agents, Ultrafast turns latency into a line item. An agent that chains dozens of model calls feels slow mostly because each call waits on token generation, so an up-to-8x speed gain can change what users will tolerate. But at six times the token price, the right question is which steps are latency-critical. Route only those to Ultrafast and keep background work on Standard or cheaper models.
CFOs should ask for cost per completed task, not cost per token. A faster model can cut developer waiting time and compute retries, but output tokens at $300 per million add up quickly in coding agents that write long files. Set per-team budgets and alerts before eligible ChatGPT Work and Codex users start defaulting to the fast tier.
Compliance and CISO teams have a narrower issue: according to OrcaRouter, Ultrafast does not support EU or other non-US regional processing endpoints. Companies with EU data residency commitments should confirm with OpenAI before sending regulated data through the tier, and check that contracts reflect where inference actually runs.
The bigger picture
Model makers are increasingly selling the same model at several speeds, with price rising with speed. The Nvidia post also served as a showcase for Blackwell as inference hardware at a time when Nvidia faces competition for inference workloads. OpenAI's comments about using its own models to tune GPU code suggest inference efficiency is becoming a competitive lever alongside model quality.
What’s next
Watch whether OpenAI raises the low launch rate limits, adds regional processing for Ultrafast, and extends the tier to other models; OrcaRouter said GPT-5.6 Sol has preview access. Buyers should benchmark Ultrafast on their own agent workloads before committing budgets.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error









