Technology

OpenAI Unveils GPT-5.6 Sol Ultrafast Mode at 14X Speed

OpenAI launches limited preview of GPT-5.6 Sol Ultrafast mode, delivering up to 750 tokens per second at 14 times standard speed on Cerebras chips.

OpenAI has launched a limited preview of Ultrafast mode for GPT-5.6 Sol, delivering up to 750 tokens per second at 14 times standard speed. The service runs on Cerebras wafer-scale chips rather than GPUs, preserving the model's full intelligence while dramatically reducing latency. Early testing spans incident response, financial research, customer support, and live research workflows.

OpenAI has opened a limited preview of Ultrafast, a new service tier for GPT-5.6 Sol that runs up to 14 times faster than Standard processing[reference:0][reference:1]. The tier generates up to 750 output tokens per second and is served on Cerebras wafer-scale hardware rather than on GPUs[reference:2][reference:3].

Same Intelligence, Radical Speed

Ultrafast mode uses the same GPT-5.6 Sol model weights and delivers identical intelligence to the standard tier[reference:4][reference:5]. The speed gain comes entirely from the underlying hardware architecture. Cerebras builds a single chip the size of a wafer, carrying 44 GB of on-chip SRAM[reference:6]. This design eliminates the bottleneck of moving model weights between separate memory and compute units, which is the limiting factor on traditional GPU clusters[reference:7].

OpenAI researcher Jeffrey Wang described the practical impact: "Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes before I even have the opportunity to context-switch"[reference:8].

Strategic Shift in AI Infrastructure

The preview marks a significant departure from OpenAI's GPU-dominated infrastructure. OpenAI is among the largest consumers of Nvidia hardware globally, but inference and training have different economics[reference:9]. Training rewards raw parallel throughput, while inference rewards latency and cost per token at enormous scale — the workload where alternative architectures have a genuine case[reference:10].

The tier went into limited preview on August 13, available to a selected group of OpenAI API customers, with access expanding as capacity allows[reference:11][reference:12].

Early Use Cases

OpenAI is testing Ultrafast across several time-sensitive applications[reference:13]:

  • Incident response: Analyzing logs, code changes, and reports during active outages[reference:14]
  • Financial research: Assessing market signals and transactions in real time[reference:15]
  • Customer support: Resolving complex issues without interrupting conversations[reference:16]
  • Commerce: Answering product questions and resolving checkout issues while shoppers are still deciding[reference:17]
  • Live research: Turning overnight experiments into interactive working sessions[reference:18]

Internally, OpenAI engineers are using Ultrafast for incident response — reading logs, analyzing traces, and synthesizing information to identify fixes in a fraction of the time[reference:19].

Availability and Pricing

No pricing has been published for the Ultrafast tier[reference:20][reference:21]. The standard GPT-5.6 Sol tier currently costs 5permillioninputtokensand5 per million input tokens and 30 per million output tokens[reference:22].

Ultrafast is not a smaller or distilled model — only the hardware and scheduling changed[reference:23]. OpenAI frames it as "more useful work per second" rather than a quality tier[reference:24]. If your business requires frontier intelligence at the highest speed, you can sign up to get notified when access expands[reference:25].