EnterpriseModelsTools

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

Source: OpenAI

Intel Summary

OpenAI has announced a preview of Ultrafast mode, a new API service tier designed for its GPT-5.6 Sol model. According to OpenAI, the tier delivers generation speeds up to 14 times faster than standard execution, reaching up to 750 output tokens per second. The service is powered by specialized AI acceleration hardware from Cerebras. The offering targets latency-sensitive applications, high-throughput enterprise data processing, and interactive agentic workflows.

Why It Matters

Inference speeds reaching 750 tokens per second materially improve the viability of real-time voice systems, complex multi-step reasoning, and low-latency software development pipelines. The partnership also marks a notable infrastructure diversification for OpenAI, demonstrating commercial deployment of Cerebras hardware alongside traditional GPU clusters in production enterprise environments.

Part of an ongoing development

Primary source

OpenAI previews Ultrafast mode for GPT-5.6 Sol

OpenAI has announced a preview of Ultrafast mode, a new API service tier designed for its GPT-5.6 Sol model. According to OpenAI, the tier delivers generation speeds up to 14 times faster than standard execution, reaching up to 750 output tokens per second. Claims are as reported; this summary makes no determination about accuracy or significance.

Confidence
Moderate confidence
Corroboration
Limited corroboration

What we know

  • Product:GPT-5.6 Sol
  • Version:5.6
  • Availability:Announced
  • Organization:OpenAI

Organizations & Entities

Topics