Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
Source: OpenAI
Intel Summary
OpenAI has announced a preview of Ultrafast mode, a new API service tier designed for its GPT-5.6 Sol model. According to OpenAI, the tier delivers generation speeds up to 14 times faster than standard execution, reaching up to 750 output tokens per second. The service is powered by specialized AI acceleration hardware from Cerebras. The offering targets latency-sensitive applications, high-throughput enterprise data processing, and interactive agentic workflows.
Why It Matters
Inference speeds reaching 750 tokens per second materially improve the viability of real-time voice systems, complex multi-step reasoning, and low-latency software development pipelines. The partnership also marks a notable infrastructure diversification for OpenAI, demonstrating commercial deployment of Cerebras hardware alongside traditional GPU clusters in production enterprise environments.
Part of an ongoing development
Primary sourceOpenAI previews Ultrafast mode for GPT-5.6 Sol
OpenAI has announced a preview of Ultrafast mode, a new API service tier designed for its GPT-5.6 Sol model. According to OpenAI, the tier delivers generation speeds up to 14 times faster than standard execution, reaching up to 750 output tokens per second. Claims are as reported; this summary makes no determination about accuracy or significance.
- Confidence
- Moderate confidence
- Corroboration
- Limited corroboration
What we know
- Product:GPT-5.6 Sol
- Version:5.6
- Availability:Announced
- Organization:OpenAI
Organizations & Entities
Topics
Related Intelligence
- DevelopmentNewAlso involving GPT-5.6
OpenAI makes GPT-5.6 available for government use via ChatGPT Enterprise
OpenAI has made its GPT-5.6 models available for government workloads through its FedRAMP-authorized ChatGPT Enterprise offering. Claims are as reported; this summary makes no determination about accuracy or significance.
1 reporting source - DevelopmentNewAlso involving GPT-5.6
OpenAI developing persistent execution feature for Codex
OpenAI is developing an experimental "Persistent Mode" for its Codex AI agent, allowing the system to run indefinitely in the background and autonomously create follow-up tasks. Code references discovered by WIRED were confirmed by OpenAI. Claims are as reported; this summary makes no determination about accuracy or significance.
1 reporting source - DevelopmentNewAlso involving GPT-5.6
AWS launches OpenAI GPT-5.6 models on Amazon Bedrock for in-country inferencing in India
Amazon Web Services has expanded Amazon Bedrock to support OpenAI GPT-5.6 models, specifically the Terra and Luna variants, for in-country inferencing within India. The deployment leverages India geographic cross-Region routing, allowing enterprises to process generative AI workloads at scale while ensuring that inference requests and associated data remain strictly within Indian territory to meet local data residency requirements. Claims are as reported; this summary makes no determination about accuracy or significance.
- ReportAlso involving GPT-5.6
Introducing cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock
AWS announced the availability of cross-Region inference for OpenAI GPT-5.6 models (Sol, Terra, and Luna) on Amazon Bedrock across more than 25 AWS Regions. The feature introduces US geographic and global inference profiles designed to dynamically route requests for increased throughput and availability. Organizations can invoke the models using either the OpenAI API or Bedrock Converse API while utilizing standard AWS Identity and Access Management (IAM), quotas, and monitoring tools.
AWS Machine Learning