OpenAI has opened a limited preview of Ultrafast, a new service tier for GPT-5.6 Sol that runs up to 14 times faster than Standard processing[reference:0][reference:1]. The tier generates up to 750 output tokens per second and is served on Cerebras wafer-scale hardware rather than on GPUs[reference:2][reference:3].
Same Intelligence, Radical Speed
Ultrafast mode uses the same GPT-5.6 Sol model weights and delivers identical intelligence to the standard tier[reference:4][reference:5]. The speed gain comes entirely from the underlying hardware architecture. Cerebras builds a single chip the size of a wafer, carrying 44 GB of on-chip SRAM[reference:6]. This design eliminates the bottleneck of moving model weights between separate memory and compute units, which is the limiting factor on traditional GPU clusters[reference:7].
OpenAI researcher Jeffrey Wang described the practical impact: "Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes before I even have the opportunity to context-switch"[reference:8].
Strategic Shift in AI Infrastructure
The preview marks a significant departure from OpenAI's GPU-dominated infrastructure. OpenAI is among the largest consumers of Nvidia hardware globally, but inference and training have different economics[reference:9]. Training rewards raw parallel throughput, while inference rewards latency and cost per token at enormous scale — the workload where alternative architectures have a genuine case[reference:10].
The tier went into limited preview on August 13, available to a selected group of OpenAI API customers, with access expanding as capacity allows[reference:11][reference:12].
Early Use Cases
OpenAI is testing Ultrafast across several time-sensitive applications[reference:13]:
- Incident response: Analyzing logs, code changes, and reports during active outages[reference:14]
- Financial research: Assessing market signals and transactions in real time[reference:15]
- Customer support: Resolving complex issues without interrupting conversations[reference:16]
- Commerce: Answering product questions and resolving checkout issues while shoppers are still deciding[reference:17]
- Live research: Turning overnight experiments into interactive working sessions[reference:18]
Internally, OpenAI engineers are using Ultrafast for incident response — reading logs, analyzing traces, and synthesizing information to identify fixes in a fraction of the time[reference:19].
Availability and Pricing
No pricing has been published for the Ultrafast tier[reference:20][reference:21]. The standard GPT-5.6 Sol tier currently costs 30 per million output tokens[reference:22].
Ultrafast is not a smaller or distilled model — only the hardware and scheduling changed[reference:23]. OpenAI frames it as "more useful work per second" rather than a quality tier[reference:24]. If your business requires frontier intelligence at the highest speed, you can sign up to get notified when access expands[reference:25].