OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 750 tokens per second

OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 750 tokens per second OpenAI has opened a preview of "Ultrafast" mode for GPT-5.6 Sol, its flagship reasoning model. The mode streams up to 750 output tokens per second, which OpenAI frames as roughly 14 times the speed of the standard API, by running inference on Cerebras hardware. Cerebras and OpenAI signed a ten billion dollar partnership earlier this year, and Ultrafast is the first major use of that capacity. Access is limited for now to

1 min
OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 750 tokens per second

OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 750 tokens per second

OpenAI has opened a preview of "Ultrafast" mode for GPT-5.6 Sol, its flagship reasoning model. The mode streams up to 750 output tokens per second, which OpenAI frames as roughly 14 times the speed of the standard API, by running inference on Cerebras hardware.

CloudSEK supply-chain breach visualization

Cerebras and OpenAI signed a ten billion dollar partnership earlier this year, and Ultrafast is the first major use of that capacity. Access is limited for now to the OpenAI API and to a set of selected customers. OpenAI says it will widen availability as capacity grows and is collecting sign-ups through a public form.

The pitch is speed without downsizing. OpenAI says Ultrafast keeps the full capabilities of a large reasoning model while matching the responsiveness of a smaller one, which it calls more useful work per second. The company points to live incident response, where logs, code changes, and postmortems could be analyzed while an outage is still unfolding, and to finance, support, and research workflows that currently run as overnight batch jobs.

Ultrafast is also a pricing move. OpenAI already sells a Fast Mode API tier that promises about 2.5x speed for GPT-5.6 Sol at roughly twice the price. Ultrafast adds a third, faster, and likely more expensive tier, turning inference latency into a metered product the way cloud providers meter compute performance.

Sources:

Written by

More to read

  • Modern Hopfield Networks: How Continuous Energy Landscapes Explain Transformer Attention and Exponential Memory

    When Vaswani et al. introduced the Transformer architecture in 2017, scaled dot-product self-attention was presented primarily as a pragmatic computational mechanism: an efficient, highly parallelizable alternative to recurrence and convolutions. By computing pairwise inner products between queries and keys, normalizing via softmax, and taking a weighted sum of values, attention allowed models to route information dynamically across arbitrarily distant tokens. For several years, self-attention

    1 min
  • Fine-Grained Access Control in Enterprise RAG: Pre-Filtering vs. Post-Filtering, Zanzibar ReBAC Models, and Zero-Trust Retrieval Architecture

    Deploying Retrieval-Augmented Generation (RAG) across enterprise knowledge repositories introduces a security boundary that simple vector search was never designed to enforce. In corporate environments spanning Google Workspace, Microsoft SharePoint, Notion, Confluence, and internal ticket systems, access permissions are dynamic, hierarchical, and deeply nested. Attempting to enforce security at the prompt generation layer by instructing language models to ignore unauthorized context is fundame

    1 min
  • Inside Ulanqab: How Inner Mongolia Became the 12.5GW Epicenter of China's AI Data Center Boom

    Located approximately 350 kilometers northwest of Beijing, the grassland municipality of Ulanqab in Inner Mongolia has transformed into China's primary hub for artificial intelligence compute infrastructure. Historically recognized for agriculture and mineral extraction, the city now hosts nearly 100 enterprise data centers operating or under active construction, with technology firms pledging an aggregate capacity of 12.5 gigawatts (GW). According to a research note published by Goldman Sachs,

    1 min