Skip to content
TECH CEO Daily
AILaunch

OpenAI GPT-6 Astra Ultrafast runs up to 8x faster on Nvidia Blackwell, at 6x the price

Nvidia says OpenAI's faster Astra tier runs on Blackwell GPUs; a pricing breakdown puts it at six times Standard token prices with only US and global processing.

By · Editor

· 3 min read · Fact-checked

The 60-second brief

  • 1Nvidia says GPT-6 Astra Ultrafast generates tokens up to 8x faster than Astra Standard mode, on Blackwell GPUs.
  • 2An OrcaRouter breakdown of OpenAI's pricing lists Ultrafast at $60 input and $300 output per million tokens, 6x Standard.
  • 3Ultrafast supports only US and global processing, not EU regional endpoints, according to OrcaRouter.

The news

OpenAI's GPT-6 Astra Ultrafast, a faster tier of the company's top model, runs on Nvidia (NVDA) Blackwell GPUs and generates tokens up to 8x faster than Astra's Standard mode, Nvidia said in a blog post on October 1, 2026. The tier targets coding, tool use and agent workflows.

According to Nvidia, Ultrafast is available in the OpenAI API and to eligible ChatGPT Work and Codex users. Nvidia pitched the tier at interactive applications and agentic AI, where a model may call tools and generate code many times before a person sees a result.

Two OpenAI executives were quoted in Nvidia's post. Uday Ruddarraju, described as OpenAI's CTO of compute, said OpenAI used its internal models to optimize inference on Nvidia GPUs. Philippe Tillet, OpenAI's inference lead, credited Nvidia's tooling and documentation for helping OpenAI's models program Blackwell and Rubin GPUs.

Nvidia's post did not list prices. A pricing breakdown published by OrcaRouter, an AI model routing service, citing OpenAI's API pricing page, lists Ultrafast at $60 per million input tokens, $6 for cached input, $75 for cache writes and $300 per million output tokens for requests up to 272,000 input tokens. That is six times the Standard tier, OrcaRouter said. Requests above 272,000 input tokens are repriced in full at $120 input and $450 output per million tokens.

OrcaRouter also reported that Ultrafast launched with low rate limits: 500,000 tokens per minute for usage tiers 1 to 3, 1 million for tier 4 and 5 million for tier 5. It said Ultrafast supports only US data residency and global processing, with no EU or other regional processing endpoints, and said Ultrafast moved from a waitlist to general availability on September 29, 2026.

The numbers

Speed vs Astra Standard (Nvidia)
Up to 8x
Price multiplier vs Standard (OrcaRouter)
6x
Input / output price, ≤272K tokens
$60 / $300 per million tokens
Input / output price, >272K tokens
$120 / $450 per million tokens
Rate limit, tiers 1-3
500,000 tokens per minute

Why CEOs should care

For engineering leaders building real-time agents, Ultrafast turns latency into a line item. An agent that chains dozens of model calls feels slow mostly because each call waits on token generation, so an up-to-8x speed gain can change what users will tolerate. But at six times the token price, the right question is which steps are latency-critical. Route only those to Ultrafast and keep background work on Standard or cheaper models.

CFOs should ask for cost per completed task, not cost per token. A faster model can cut developer waiting time and compute retries, but output tokens at $300 per million add up quickly in coding agents that write long files. Set per-team budgets and alerts before eligible ChatGPT Work and Codex users start defaulting to the fast tier.

Compliance and CISO teams have a narrower issue: according to OrcaRouter, Ultrafast does not support EU or other non-US regional processing endpoints. Companies with EU data residency commitments should confirm with OpenAI before sending regulated data through the tier, and check that contracts reflect where inference actually runs.

The bigger picture

Model makers are increasingly selling the same model at several speeds, with price rising with speed. The Nvidia post also served as a showcase for Blackwell as inference hardware at a time when Nvidia faces competition for inference workloads. OpenAI's comments about using its own models to tune GPU code suggest inference efficiency is becoming a competitive lever alongside model quality.

What’s next

Watch whether OpenAI raises the low launch rate limits, adds regional processing for Ultrafast, and extends the tier to other models; OrcaRouter said GPT-5.6 Sol has preview access. Buyers should benchmark Ultrafast on their own agent workloads before committing budgets.

What “Fact-checked” means

Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.

What we checked
Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
How
A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
Who
The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, . A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
If something is wrong
“Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error

How we fact-check →

Companies in this story

Earlier coverage of OpenAI

All OpenAI coverage →

Written by

Editor · Technology & Business Writer

Hussein is a writer and business technology enthusiast focused on the intersection of technology, entrepreneurship, finance, artificial intelligence, and digital innovation.

CoversAICybersecurityBig TechSaaSStartupsFintech

About this story. Researched from primary sources whenever they are available and fact-checked before publication.

Published by Tech CEO Daily, an independent publication. Masthead · Editorial standards

Follow Tech CEO Daily on Facebook for the day’s top stories in your feed.

Free newsletters

The technology briefing for people running businesses.

Daily, weekly, bi-weekly or monthly. You choose.

How often

The Daily Brief · Monday to Saturday, 7 a.m. ET

Free forever. One click to unsubscribe. We never sell your email.