ULTRAFAST
53 tok/s
scroll = throttle

OpenAI announcement explainer · powered by Cerebras

Ultrafast: GPT‑5.6 Sol,
14× the speed.

A new service tier in the OpenAI API runs OpenAI's most intelligent model at up to 750 output tokens per second — up to 14 times faster than Standard processing. Until now, real-time speed usually meant choosing a smaller or more specialized model. Ultrafast points the other way: more useful work per second.

First in the OpenAI API Up to 750 tok/s Powered by Cerebras Limited preview

01 · The headline numbers

Four numbers to remember

Everything else on this page hangs off these. They count up as they arrive — like everything else here, they don't make you wait.

0×
faster than Standard processing (upper bound)
0
output tokens per second, peak
GPT-5.6 Sol
the most intelligent model, unchanged
API1st
launches first in the OpenAI API

Figures from OpenAI's announcement, Aug 13 2026 · “Previewing Ultrafast mode.” Both headline rates are “up to” peaks; your mileage will vary.

02 · The speed duel

Same prompt. Same model. One of them is waiting for you.

In the announcement's demo, Ultrafast and Standard built the same working 3D warehouse simulator side by side from one text prompt. Here's that race in miniature — fire it yourself.

Both lanes get the identical 200-character brief.
STANDARD · GPT‑5.6 Sol ≈54 tok/s
ULTRAFAST · GPT‑5.6 Sol 750 tok/s

Simulation at announcement rates (750 vs ≈53.6 tok/s derived from 14×). The original demo generated a 3D warehouse simulator — no warehouses were harmed in this browser.

03 · The silicon story

Why wafer-scale silicon moves this fast

Ultrafast is powered by Cerebras. Their Wafer‑Scale Engine doesn't shuttle data across boards and cables — the memory lives next to every core on one enormous chip. Fewer round trips, less latency.

One wafer, one chip. Every square has compute plus its own SRAM — weights are read where they live. (concept diagram)
largest GPU 814 mm² WSE‑3 ≈ 46,225 mm² · 57× larger

WSE‑3, by the numbers

  • Silicon area46,225 mm² 57× the largest GPU (814 mm²)
  • Transistors4 trillion vs ~80 billion in the leading GPU
  • AI cores900,000 84 dies · ≈10,700 cores each
  • On-chip memory44 GB SRAM single-clock-cycle access from every core
  • Memory bandwidth21 PB/s ≈7,000× the leading GPU (Cerebras)
  • On-wafer fabric214 Pb/s no board traces, no cables between cores

Hardware figures: Cerebras WSE‑3 datasheet & HotChips 2024 presentation, cerebras.ai — not announcement claims. Ultrafast itself is characterized only by the canonical pair: powered by Cerebras, up to 750 tok/s.

04 · Latency economics

What 14× does to a workday

Generation time ≈ tokens ÷ rate. Drag the slider to size a task, then watch the wall clock collapse. Illustrative arithmetic on the announcement's two rates.

typical incident brief ≈ 1–3k tokens
14× less waiting
standard ≈53.6 tok/s
ultrafast 750 tok/s
🌙 Overnight research batch overnight → hours → multiple loops before lunch research iterations tighten from overnight runs into the workday (announcement)
🚨 Incident diagnosis minutes of waiting → fits inside the outage window logs + code changes + engineer reports analyzed while the outage unfolds
💬 Customer conversation awkward pauses → resolves mid-conversation complex issues handled in real time, across steps and systems
🛒 Checkout hesitation answers after doubt → answers while they browse product questions & inventory answered before hesitation becomes an abandoned cart

05 · Use-case gallery

Where every second earns its keep

Five scenarios from the announcement. Hover, tap, or tab through them — each carries the moment where latency decides the outcome.

✓ EXPLORED

Incident response & reliability

Analyze logs, recent code changes, and engineer reports to find the likely cause — and prepare a fix — while the outage is still unfolding.

Seconds decide whether the fix lands inside the outage window
✓ EXPLORED

Financial research & security

Analyze market signals, assess transactions, identify suspicious activity while conditions are still changing.

Markets reprice in milliseconds; analysis must keep pace
✓ EXPLORED

Customer support & voice

Resolve complex issues in real time without interrupting the conversation — even across multiple steps or systems.

Silence on the line is where conversations die
✓ EXPLORED

Commerce

Answer product questions, check inventory, personalize recommendations, resolve checkout issues — before hesitation becomes an abandoned cart.

Every pause at checkout is a door quietly closing
✓ EXPLORED

Live research & experimentation

What used to be an overnight batch run becomes an interactive working session.

Ideas iterate at the speed of curiosity, not the speed of queues

06 · Proof from the field

Early hands on the throttle

The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.

John Crepezzi, AI Assistants, Jane Street · featured early customer
JANE STREETEARLY CUSTOMER
PODIUMEARLY CUSTOMER
BASISEARLY CUSTOMER
ROGOEARLY CUSTOMER
LIMITED PREVIEW · LIVE NOW

You've reached 750.

GPT‑5.6 Sol on Ultrafast is in a limited preview today for a select group of customers. Access expands as capacity grows — sign up on OpenAI's Ultrafast page to be notified when it widens.

Links leave this offline artifact when clicked — the page itself never phones home.

Get up to 40% off Z.ai