Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

2026-08-20 · OpenAI

Preview Ultrafast Mode: GPT-5.6 Sol Speed Boosted up to 14 Times

Release Overview

OpenAI has released a new API service tier named Preview Ultrafast. This tier is built to deliver substantially faster generation performance for the GPT-5.6 Sol model.

The announcement centers on significant speed improvements achieved through specialized hardware acceleration.

Technical Foundation

The Preview Ultrafast tier is based on Cerebras hardware acceleration. This forms the technical basis for the performance gains.

By leveraging Cerebras hardware, OpenAI has enabled a major increase in the model's output speed.

Performance Details

The generation speed of the GPT-5.6 Sol model can be boosted by up to 14 times using this tier. This multiplier represents the primary performance claim.

Real-world testing demonstrates an output rate of approximately 750 tokens per second. This measured speed is a key highlighted metric.

Key Metrics

  • Model: GPT-5.6 Sol
  • Maximum Speed Increase: 14x
  • Measured Output: ~750 tokens per second

Targeted Application Scenarios

The new tier aims to meet the needs of applications with strict low-latency requirements.

Specific examples provided include:

  • Real-time conversation
  • Interactive content generation

These scenarios demand minimal delay between input and output. The Preview Ultrafast tier is designed specifically to address such strict latency demands by delivering much faster token generation.

Benefits for Developers

The service provides developers with a significantly faster response experience.

When using the Preview Ultrafast tier, developers will encounter reduced latency when generating content with the GPT-5.6 Sol model. This faster response enables improved application performance in time-sensitive projects.

Future Optimization Plans

OpenAI stated that it will further optimize the throughput capability of this tier in the future.

This statement indicates ongoing work to enhance the tier's overall capacity beyond the current performance levels.

Comprehensive Summary

OpenAI's Preview Ultrafast API service tier, powered by Cerebras hardware acceleration, increases the generation speed of the GPT-5.6 Sol model by up to 14 times, with real-world tests showing approximately 750 tokens per second. The tier is intended for applications with demanding low-latency needs, such as real-time conversation and interactive content generation, while offering developers faster response times.

OpenAI has also indicated plans to further optimize the throughput of the Preview Ultrafast tier going forward. All details in the announcement focus exclusively on the speed improvement, the hardware basis, the targeted low-latency use cases, the benefits to developers, and the commitment to future throughput optimization.

The structured rollout emphasizes that this tier is purpose-built for scenarios where response speed is critical. By achieving up to 14x faster generation and 750 tokens per second, it directly addresses the requirements of real-time dialogue systems and interactive tools. Developers gain a practical way to reduce latency in their applications, and OpenAI's stated intention to improve throughput suggests continued refinement of the service.

This release contains no information beyond the speed multiplier, token rate, hardware foundation, listed use cases, developer benefits, and future optimization commitment.

Source