OpenAI Runs GPT-5.6 Sol At 750 Tokens/Second On Cerebras In New Ultrafast Tier

Cerebras is powering OpenAI's new Ultrafast API tier, running GPT-5.6 Sol at up to 750 output tokens per second and 14x faster than standard processing.

OpenAI Runs GPT-5.6 Sol At 750 Tokens/Second On Cerebras In New Ultrafast Tier

Cerebras (NASDAQ: CBRS) will power OpenAI's new Ultrafast mode, a service tier that runs the flagship GPT-5.6 Sol model at up to 750 output tokens per second and up to 14 times faster than the standard OpenAI API. Announced August 13 from Sunnyvale, the limited-preview offering targets latency-critical workloads — financial research, incident response, voice apps, commerce, and live experimentation — where token throughput determines whether AI can sit inside the interaction loop.

Frontier Intelligence At Broadband Speed

“GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive,” said Cerebras CEO Andrew Feldman. OpenAI VP Sachin Katti described the launch as an exploration of “what becomes possible when customers can get the intelligence of our most capable models with significantly lower latency.” Ultrafast preserves the intelligence of standard GPT-5.6 Sol; it is not a distilled or smaller variant.

Benchmarks: 7x Faster On Humanity's Last Exam

On Humanity's Last Exam, a 2,500-question graduate-level benchmark, GPT-5.6 Sol Ultrafast finished in just over 11 hours — against more than three days for Claude Fable 5 at comparable accuracy, roughly 7x faster. On GDP-Val, a benchmark of economically valuable knowledge-work tasks like legal briefs and financial models, Ultrafast delivered a 5.6x end-to-end speedup with no quality loss.

OpenAI logo

Wafer-Scale Memory Bypasses The GPU Bottleneck

Ultrafast's speed comes from Cerebras' Wafer-Scale Engine, which keeps 44 GB of model weights on 44 GB of SRAM per wafer-sized chip rather than shuttling them to and from off-chip HBM as GPU inference does. That eliminates the memory-bandwidth bottleneck that caps frontier-model inference speed on Nvidia hardware. Related coverage: Anthropic's in-house chip effort and xAI Grok 4.6 launch.

Reporting based on coverage from Cerebras Investor Relations, OpenAI and Unite.AI.

Category: Edge Computing

Tags: AI Enterprise AI AI Foundation Models OpenAI AI Chips

Related Articles