Cerebras stock slides as Nvidia reportedly powers OpenAI’s ’Ultrafast’ tier
Investing.com -- Shares of Cerebras Systems tumbled 7% following industry speculation that OpenAI bypassed the startup’s specialized AI hardware to power its newly announced high-speed AI tier, GPT-6.1 Sol Ultrafast.
The sell-off was triggered when semiconductor research firm SemiAnalysis posted on X (formerly Twitter) that OpenAI’s newest flagship model is running on traditional Nvidia GPUs at a low batch size, rather than on Cerebras chips. The revelation prompted immediate questions from analysts about whether Cerebras will secure a role in serving the highly anticipated GPT-6.1 Sol Ultrafast model in the future, dealing a blow to near-term investor confidence.
OpenAI recently introduced its "Ultrafast" tier, promising token generation speeds up to eight times faster than standard deployments (reaching an estimated 300 tokens per second). Generating text at these extreme speeds requires massive memory bandwidth to process requests instantly as they arrive—a computational challenge that Cerebras’ hardware was expressly designed to solve.
By positioning itself as the premier alternative to traditional GPU clusters, Cerebras has built its reputation on its Wafer-Scale Engine (WSE). Unlike standard processors, the WSE is a massive, plate-sized chip that houses billions of cores and vast amounts of memory on a single piece of silicon. This architecture allows it to hold entire massive neural networks directly on the chip, theoretically enabling lightning-fast responses without the latency of shuffling data between hundreds of separate GPUs.
The technical crux of the SemiAnalysis report lies in the phrase "low batch size".
In AI inference, "batching" refers to grouping multiple user queries together and processing them simultaneously to maximize GPU efficiency. Traditional Nvidia GPUs typically require large batch sizes to achieve peak utilization and cost-effectiveness. Conversely, a "low batch size" (processing one or a few requests at a time) prioritizes immediate response speed (low latency) over maximum throughput—an area where Cerebras’ massive on-chip memory usually holds a distinct architectural advantage.
If OpenAI has successfully optimized Nvidia GPUs to run a frontier model like GPT-6.1 Sol at a low batch size with cost-effective efficiency, it directly challenges one of Cerebras’ most potent selling points.
The speculation strikes at the heart of Cerebras’ growth narrative. Serving advanced, high-speed AI inference workloads for top-tier players like OpenAI is considered the holy grail for AI hardware challengers. While the AI chip market is expanding rapidly, Nvidia’s deeply entrenched CUDA software ecosystem and continuous hardware improvements make it a formidable incumbent.
If Nvidia’s existing infrastructure can natively achieve the ultra-low latency required for next-generation, real-time AI agents, investors fear that specialized alternatives like Cerebras may face a steeper uphill battle in securing enterprise data center contracts. Cerebras has yet to issue a public statement confirming or denying its involvement in OpenAI’s future infrastructure roadmap.
You May Also Be Interested In
- Peoples Bancorp falls, Capital Bancorp rises amid all-stock deal
- Oracle shares slip on unconfirmed report of delays at Wisconsin AI mega-campus
- Why are oil futures still so high if Middle East exports have recovered?
Create E-mail Alert Related Categories
General News, InvestingSign up for StreetInsider Free!
Receive full access to all new and archived articles, unlimited portfolio tracking, e-mail alerts, custom newswires and RSS feeds - and more!



Tweet
Share