OpenAI announcement explainer · powered by Cerebras
Ultrafast: GPT‑5.6 Sol,
14× the speed.
A new service tier in the OpenAI API runs OpenAI's most intelligent model at up to 750 output tokens per second — up to 14 times faster than Standard processing. Until now, real-time speed usually meant choosing a smaller or more specialized model. Ultrafast points the other way: more useful work per second.
Live output rate
announcement figuresThe gap you're watching is the announcement's headline: same intelligent model, radically different serving hardware.
01 · The headline numbers
Four numbers to remember
Everything else on this page hangs off these. They count up as they arrive — like everything else here, they don't make you wait.
Figures from OpenAI's announcement, Aug 13 2026 · “Previewing Ultrafast mode.” Both headline rates are “up to” peaks; your mileage will vary.
02 · The speed duel
Same prompt. Same model. One of them is waiting for you.
In the announcement's demo, Ultrafast and Standard built the same working 3D warehouse simulator side by side from one text prompt. Here's that race in miniature — fire it yourself.
Simulation at announcement rates (750 vs ≈53.6 tok/s derived from 14×). The original demo generated a 3D warehouse simulator — no warehouses were harmed in this browser.
03 · The silicon story
Why wafer-scale silicon moves this fast
Ultrafast is powered by Cerebras. Their Wafer‑Scale Engine doesn't shuttle data across boards and cables — the memory lives next to every core on one enormous chip. Fewer round trips, less latency.
WSE‑3, by the numbers
- Silicon area46,225 mm² 57× the largest GPU (814 mm²)
- Transistors4 trillion vs ~80 billion in the leading GPU
- AI cores900,000 84 dies · ≈10,700 cores each
- On-chip memory44 GB SRAM single-clock-cycle access from every core
- Memory bandwidth21 PB/s ≈7,000× the leading GPU (Cerebras)
- On-wafer fabric214 Pb/s no board traces, no cables between cores
Hardware figures: Cerebras WSE‑3 datasheet & HotChips 2024 presentation, cerebras.ai — not announcement claims. Ultrafast itself is characterized only by the canonical pair: powered by Cerebras, up to 750 tok/s.
04 · Latency economics
What 14× does to a workday
Generation time ≈ tokens ÷ rate. Drag the slider to size a task, then watch the wall clock collapse. Illustrative arithmetic on the announcement's two rates.
05 · Use-case gallery
Where every second earns its keep
Five scenarios from the announcement. Hover, tap, or tab through them — each carries the moment where latency decides the outcome.
Incident response & reliability
Analyze logs, recent code changes, and engineer reports to find the likely cause — and prepare a fix — while the outage is still unfolding.
Seconds decide whether the fix lands inside the outage windowFinancial research & security
Analyze market signals, assess transactions, identify suspicious activity while conditions are still changing.
Markets reprice in milliseconds; analysis must keep paceCustomer support & voice
Resolve complex issues in real time without interrupting the conversation — even across multiple steps or systems.
Silence on the line is where conversations dieCommerce
Answer product questions, check inventory, personalize recommendations, resolve checkout issues — before hesitation becomes an abandoned cart.
Every pause at checkout is a door quietly closingLive research & experimentation
What used to be an overnight batch run becomes an interactive working session.
Ideas iterate at the speed of curiosity, not the speed of queues06 · Proof from the field
Early hands on the throttle
The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.
You've reached 750.
GPT‑5.6 Sol on Ultrafast is in a limited preview today for a select group of customers. Access expands as capacity grows — sign up on OpenAI's Ultrafast page to be notified when it widens.
Links leave this offline artifact when clicked — the page itself never phones home.