Add MobiGyaan as a preferred source on Google
OpenAI has previewed Ultrafast, a new service tier designed to deliver much faster responses from its flagship GPT-5.6 Sol model. The offering is powered by Cerebras and is being introduced initially through the OpenAI API for a limited group of customers.
OpenAI says Ultrafast can generate up to 750 output tokens per second, significantly reducing response latency for applications where AI needs to respond in near real time. The company is testing the technology across workflows including coding, financial research, customer support, commerce, incident response, and live experimentation.

GPT-5.6 Sol Ultrafast
Ultrafast is aimed at workloads where waiting for a conventional AI response can interrupt the workflow. By combining GPT-5.6 Sol with Cerebras infrastructure, OpenAI is targeting substantially faster inference while retaining the capabilities of its flagship model.
OpenAI is testing Ultrafast for several use cases:
- Incident Response: Analyze logs, code changes, traces, and engineering reports to identify potential causes during an outage
- Financial Research & Security: Analyze changing market signals and transactions and identify potentially suspicious activity
- Customer Support & Voice: Resolve complex customer issues requiring multiple steps or information sources in real time
- Commerce: Answer product questions, check inventory, personalize recommendations, and assist with checkout
- Live Research: Run experiments, analyze results, modify approaches, and begin subsequent iterations without long waiting periods
Early Customer Testing
OpenAI is currently testing GPT-5.6 Sol with Ultrafast with a selected group of companies working across coding, commerce, financial research, customer support, and other interactive applications.
The testing is taking place in real production environments to understand how significantly faster inference changes the way businesses build and operate AI-powered products.
OpenAI’s Internal Use
OpenAI is also experimenting with Ultrafast internally, particularly for incident response and research.
During an incident, engineers can use GPT-5.6 Sol to analyze application logs and traces, summarize relevant conversations, identify additional checks, and help prepare or validate potential fixes. Engineers remain responsible for evaluating the recommendations and deploying any changes.
For research workflows, OpenAI uses the faster model to search knowledge sources, query data, and organize and summarize information across connected tools. The increased speed allows researchers to complete more experiment-and-review cycles within the same workday.
Availability
GPT-5.6 Sol with Ultrafast is currently available as a limited preview to select customers through the OpenAI API. OpenAI plans to expand access as additional capacity becomes available.
OpenAI previously announced that GPT-5.6 Sol would run on Cerebras at speeds of up to 750 tokens per second, with initial access restricted to selected customers while capacity is expanded.
