Back to today's topics

Verified · Aug 14, 2026

Independently verified

OpenAI + Cerebras: GPT-5.6 Sol Ultrafast mode on the OpenAI API — 750 t/s, 11x faster than Fable 5

2 sources

On 2026-08-13 Cerebras and OpenAI launched GPT-5.6 Sol 'Ultrafast Mode' on the OpenAI API, powered by Cerebras Wafer-Scale Engine hardware (44 GB SRAM per wafer-sized chip). Cerebras reports up to 750 output tokens/sec, 11x faster than Fable 5, 5x faster than Opus 4.8 on Fast mode, and 5.6x end-to-end speedup on GDP-Val with no quality degradation. The release completed all 2,500 Humanity's Last Exam questions in 11 hours and 11 minutes vs 78 hours 27 minutes for Claude Fable 5 (~7x faster, comparable accuracy). The mode is in limited preview, select customers only; pricing was not mentioned in the captured summary. The takeaway for creators: GPT-5.6 Sol on Cerebras is the speed-leverage bet inside the OpenAI API surface — the speed is real (Cerebras publishes on-chip SRAM as the differentiator), but limited-preview access and undisclosed pricing keep it as a 'watch this tier' story rather than a 'migrate today' story.

Why now

The story is the editorial frame for 'OpenAI + Cerebras partner for an Ultrafast tier' — useful because creators watching the speed-vs-cost frontier now have a Cerebras-hosted OpenAI API tier as a comparison point.

Why it is worth publishing

Demo potential: a 4-minute explainer on what the Cerebras Wafer-Scale Engine differentiator is, what the reported speed numbers mean for time-sensitive creator workflows, and what the limited-preview access implies for migration timing.

Evidence basis

Cerebras blog post + HN front-page coverage

OpenAI just shipped an Ultrafast tier powered by Cerebras — 750 output tokens per second, 11x faster than Fable 5 — and the differentiator is 44 GB of on-chip SRAM per wafer.

Angle

Frame the release as the speed-leverage bet inside the OpenAI API surface — use the Cerebras on-chip SRAM as the differentiator, and frame the limited-preview status as a 'watch this tier' framing rather than a 'migrate today' framing.

Format

Short talking-head video

Demo idea

A 4-minute explainer: 1 min 'what Cerebras on-chip SRAM means for token throughput', 1 min 'the reported speed numbers (750 t/s, 11x Fable 5, 5.6x GDP-Val)', 1 min 'Humanity's Last Exam 11h 11m vs 78h 27m — what 7x end-to-end speedup looks like in practice', 1 min 'limited-preview status and what to watch before migration'.

Platform notes

Pricing was not mentioned in the captured Cerebras blog summary. Frame the release as 'limited preview, select customers only' and do not quote a per-token price without re-extracting from Cerebras or OpenAI. The Cerebras page itself notes performance comparisons are based on third-party benchmarking or internal testing; real-world speed vs GPU systems may vary.

Usable claims

  • Cerebras and OpenAI launched GPT-5.6 Sol 'Ultrafast Mode' on the OpenAI API on 2026-08-13, powered by Cerebras Wafer-Scale Engine hardware (44 GB SRAM per wafer-sized chip). Cerebras reports up to 750 output tokens/sec, 11x faster than Fable 5, and 5.6x end-to-end speedup on GDP-Val with no quality degradation. Limited preview, select customers only.

Evidence pipeline

Breakdown

GPT-5.6 Sol Ultrafast mode on the OpenAI API is the Cerebras-hosted speed-leverage tier — useful to frame as the speed bet but easy to misread as a general-availability migration target. The risk: the mode is limited preview, select customers only, with pricing not mentioned in the captured summary. This explainer uses the Cerebras on-chip SRAM as the differentiator and frames the limited-preview status as a 'watch this tier' framing rather than a 'migrate today' framing.

Risks

  • Frame the release as 'limited preview, Cerebras-hosted OpenAI API tier' and flag pricing as not-on-the-record. Re-extract pricing from Cerebras or OpenAI before stating any per-token cost.

Demo ideas

  • On-chip vs off-chip framing card: '44 GB SRAM per wafer' vs GPU HBM (typically 80 GB per device, off-chip) — what this means for token-throughput ceilings.
  • Workflow comparison: same agent task on the standard OpenAI API tier vs the Ultrafast tier (if accessible) — measure latency and reliability.
  • Watch-list card: 'what to check before considering Ultrafast migration' (pricing, GA timeline, throughput stability across model loads).