Provider Dossier · Specialist Hardware
Cerebras Inference
Extreme throughput from wafer-scale engines.
The Datum Index verdict
Recommended with confidence
Ranked No. 12 of 20 in our composite reputation index, and currently improving.
≈ 4.1 / 5 aggregate · 190 reviewer notes
Composite index · reviewed 10 July 2026
Findings
- Cerebras Inference holds a Datum Index composite reputation of 82 / 100 (≈ 4.1 / 5), ranked No. 12 of 20.
- It is a specialist hardware based in Sunnyvale, USA, founded 2015, offering open-weight model access.
- Its flagship offering is Wafer-scale served open models, exposed through an OpenAI-compatible API.
- Indicative pricing is $0.60 / 1M tokens in / $0.80 / 1M tokens out (medium confidence).
- Reference throughput is ~1794 tok/s with ~530 ms time-to-first-token; stated uptime is 99.5%.
Cerebras Inference enters our index at No. 12, a placement that reads as recommended with confidence. The composite rests on 190 reviewer notes gathered across the public record, and it is best understood not as a score out of ten but as a standing among peers. The trend line is pointed upward: reviewer sentiment has strengthened over recent quarters.
In its favour
Extreme throughput from wafer-scale engines.
On the numbers that flatter it: open-weight access, 1 named compliance attestation, and a reference throughput of ~1794 tokens per second.
Points of concern
Newer public cloud; limited model and region coverage.
The caveat a buyer should price in before committing production traffic. As with every entry, we weigh it against the field rather than against perfection.
Incident & Uptime Note
The record of being there
Cerebras Inference states an availability of 99.5%. Held over a full year, that headline implies on the order of 43.8 hours of cumulative unavailability — a useful sense of scale, though real incidents cluster rather than spread evenly, and a single bad afternoon can outweigh a quiet quarter.
Observatory Panel
Latency & throughput, in distribution
Monthly panels across the models Cerebras Inference serves medium
gpt-oss-120b
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~590 | ~510 | ~1100 | ~1750 | ~1880 | ~2220 |
| June 2026 | ~610 | ~560 | ~980 | ~1580 | ~1710 | ~1890 |
| July 2026 | ~650 | ~570 | ~1060 | ~1580 | ~1680 | ~1940 |
Gemma 4 31B
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~1040 | ~930 | ~1900 | ~1190 | ~1220 | ~1430 |
| June 2026 | ~1090 | ~950 | ~1820 | ~1110 | ~1190 | ~1360 |
| July 2026 | ~1030 | ~960 | ~1860 | ~1130 | ~1170 | ~1390 |
GLM-4.7
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~790 | ~670 | ~1220 | ~920 | ~970 | ~1090 |
| June 2026 | ~690 | ~650 | ~1210 | ~970 | ~1000 | ~1160 |
| July 2026 | ~760 | ~650 | ~1250 | ~940 | ~1000 | ~1190 |
All figures are single-stream, per-request rates as one user would see them — never aggregate accelerator throughput. Serving modes are not comparable with one another: shared serverless APIs batch many tenants per accelerator, dedicated and self-served deployments hand one tenant the whole card, and specialist silicon is a regime of its own. Read each figure within its mode.
No field returns on record for Cerebras Inference yet — the ledger above opens as soon as the first practitioner submission clears verification.
Key Facts
The record, in brief
| Category | Specialist Hardware |
|---|---|
| Headquarters | Sunnyvale, USA |
| Founded | 2015 |
| Model access | Open-weight |
| Flagship / reference | Wafer-scale served open models · Llama 3.3 70B |
| OpenAI-compatible API | Yes |
| Indicative price (in / out) | $0.60 / 1M tokens / $0.80 / 1M tokens medium |
| Reference throughput | ~1794 tok/s |
| Reference TTFT | ~530 ms |
| Stated uptime | 99.5% |
| Compliance | SOC 2 Type II |
| Composite reputation | 82 / 100 · ≈ 4.1 / 5 · 190 notes |
Reviewer Notes
What the field says
The 190 notes behind Cerebras Inference's standing are drawn from the accumulated public verdict of practitioners — the recurring praises and the recurring gripes — rather than from any single survey. Two themes dominate. Admirers return to one point above all: Extreme throughput from wafer-scale engines. Detractors return to another: Newer public cloud; limited model and region coverage.
We publish neither individual reviews nor reviewer identities; the composite is an editorial synthesis, and the reviewer-note count is an order-of-magnitude indication of how much public signal informs it.
Related dossiers · see also
Peer entries in the same or adjacent category, for comparison against Cerebras Inference's standing.