Upgrade to SI Premium - Free Trial

Cerebras to power OpenAI's GPT-5.6 Sol at up to 750 tokens per second

August 13, 2026 1:00 PM

Cerebras (NASDAQ: CBRS) announced it is powering a new service tier called Ultrafast for OpenAI's GPT-5.6 Sol model, delivering output speeds of up to 750 tokens per second, which the company claims is up to 14 times faster than standard processing.

The Ultrafast tier is initially available in limited preview to OpenAI API customers. According to the press release, the service runs GPT-5.6 Sol at the same intelligence level as the standard version of the model.

Cerebras attributes the speed gains to its Wafer-Scale Engine architecture, which stores 44 GB of SRAM on each wafer-sized chip, keeping model weights on-chip rather than transferring data between on-chip and off-chip memory as is required in GPU-based inference systems.

The companies cited benchmark results to support the speed claims. On Humanity's Last Exam, a 2,500-question graduate-level test, GPT-5.6 Sol Ultrafast completed the full question set in just over 11 hours, compared to more than three days for Claude Fable 5, with comparable accuracy. On GDP-Val, a benchmark of knowledge-work tasks including legal briefs and financial models, Ultrafast delivered a 5.6 times end-to-end speedup with no reported loss in quality.

Sachin Katti, VP Compute Strategy and GPT-Infra at OpenAI, said the company is "starting with a small group of customers to learn where that speed creates meaningful value" before expanding the service.

Andrew Feldman, CEO and co-founder of Cerebras, said the partnership demonstrates that "speed and intelligence are no longer mutually exclusive."

Categories

Corporate News

Next Articles