Trending: On-device modelsSearch
iHeartGeek
iTECH

NVIDIA's Blackwell GPUs accelerate GPT-6 Astra Ultrafast

OpenAI's Ultrafast tier for GPT-6 Astra runs on NVIDIA Blackwell hardware and promises up to eight times the token generation rate of the standard mode, the chipmaker says.

A comparison graphic showing a Standard mode at 17 seconds beside an Ultrafast mode at 6.4 seconds, drawn as two rocket illustrations with bars charting response time

The fast tier of OpenAI's GPT-6 Astra model runs on NVIDIA Blackwell GPUs, with up to eight times the token generation rate of the standard mode, according to NVIDIA.

A speed tier rather than a new model

GPT-6 Astra Ultrafast is available now through the OpenAI API and to eligible ChatGPT Work and Codex users. NVIDIA frames the gain around repeated work rather than single questions: an agent writes code, calls a tool, checks the result and decides what to do next, and every one of those steps waits on a model response. Faster generation shortens the edit-test-debug loop, cuts the dead time between tool calls and is the difference between an interactive application that feels responsive and one that does not.

Optimisation that continues after deployment

NVIDIA says the speed comes from continuous inference optimisation as much as from the silicon, and that OpenAI keeps refining its serving software after a model ships rather than treating launch as the end of the work. Philippe Tillet, OpenAI's inference lead, is quoted saying NVIDIA's investment in tooling and documentation has made OpenAI's models unusually good at writing software for Blackwell and Rubin GPUs, and that Astra turns that knowledge into high-performance kernels spanning latency, throughput and cost.

Where the infrastructure argument sits

The post is a useful marker of how the competition between model providers has moved. Model quality still decides which assistant a developer opens, but the price and latency of serving that model decide whether it can sit inside a loop that runs hundreds of times a day — and that makes the accelerator underneath a first-order concern rather than a detail. NVIDIA's interest in the framing is plain, since the faster tier exists because Blackwell does the work, and the figures come from the company's own blog rather than an independent benchmark. OpenAI's newsroom could not be read from this machine, so no OpenAI link accompanies the figures.

Our opinion

There is something quietly telling about the fact that the interesting number in a flagship AI release is now a latency figure with a decimal point. “Eight times faster” sounds like a marketing brag until you notice what it is measured against: a standard mode that already had to be usable, which means the comparison is between fast and faster rather than between working and broken. That is what a maturing platform looks like — the headline is no longer capability but the cost of using it constantly. It also puts the squeeze exactly where developers feel it. An agent that takes seventeen seconds to answer is a novelty; an agent that takes six is something you leave running in the background while you do other work, and that change in habit is worth more than any benchmark table.