Provider Dossier · Aggregator
Replicate
Broadest catalog including image, audio and video models.
The Datum Index verdict
Recommended with confidence
Ranked No. 15 of 20 in our composite reputation index, and currently holding steady.
≈ 4.0 / 5 aggregate · 410 reviewer notes
Composite index · reviewed 10 July 2026
Findings
- Replicate holds a Datum Index composite reputation of 80 / 100 (≈ 4.0 / 5), ranked No. 15 of 20.
- It is a aggregator based in San Francisco, USA, founded 2019, offering open-weight model access.
- Its flagship offering is Community + language models, through its own API.
- Indicative pricing is $0.65 / 1M tokens in / $2.75 / 1M tokens out (medium confidence).
- Reference throughput is ~50 tok/s with ~2500 ms time-to-first-token; stated uptime is 99.3%.
Replicate enters our index at No. 15, a placement that reads as recommended with confidence. The composite rests on 410 reviewer notes gathered across the public record, and it is best understood not as a score out of ten but as a standing among peers. The trend line is flat: reviewer sentiment has held steady over recent quarters.
In its favour
Broadest catalog including image, audio and video models.
On the numbers that flatter it: open-weight access, 1 named compliance attestation, and a reference throughput of ~50 tokens per second.
Points of concern
Cold starts and its own (non-OpenAI) API shape for LLMs.
The caveat a buyer should price in before committing production traffic. As with every entry, we weigh it against the field rather than against perfection.
Incident & Uptime Note
The record of being there
Replicate states an availability of 99.3%. Held over a full year, that headline implies on the order of 61.3 hours of cumulative unavailability — a useful sense of scale, though real incidents cluster rather than spread evenly, and a single bad afternoon can outweigh a quiet quarter.
Observatory Panel
Latency & throughput, in distribution
Monthly panels across the models Replicate serves medium
Llama 3.3 70B
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~2710 | ~2520 | ~4760 | ~48 | ~50 | ~56 |
| June 2026 | ~2820 | ~2460 | ~4340 | ~48 | ~51 | ~58 |
| July 2026 | ~2650 | ~2440 | ~5360 | ~48 | ~51 | ~60 |
DeepSeek V4 Flash
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~2920 | ~2560 | ~4020 | ~52 | ~56 | ~60 |
| June 2026 | ~3110 | ~2800 | ~4830 | ~50 | ~51 | ~55 |
| July 2026 | ~2880 | ~2480 | ~4040 | ~55 | ~58 | ~68 |
All figures are single-stream, per-request rates as one user would see them — never aggregate accelerator throughput. Serving modes are not comparable with one another: shared serverless APIs batch many tenants per accelerator, dedicated and self-served deployments hand one tenant the whole card, and specialist silicon is a regime of its own. Read each figure within its mode.
No field returns on record for Replicate yet — the ledger above opens as soon as the first practitioner submission clears verification.
Key Facts
The record, in brief
| Category | Aggregator |
|---|---|
| Headquarters | San Francisco, USA |
| Founded | 2019 |
| Model access | Open-weight |
| Flagship / reference | Community + language models · Llama 3.3 70B |
| OpenAI-compatible API | No |
| Indicative price (in / out) | $0.65 / 1M tokens / $2.75 / 1M tokens medium |
| Reference throughput | ~50 tok/s |
| Reference TTFT | ~2500 ms |
| Stated uptime | 99.3% |
| Compliance | SOC 2 Type II |
| Composite reputation | 80 / 100 · ≈ 4.0 / 5 · 410 notes |
Reviewer Notes
What the field says
The 410 notes behind Replicate's standing are drawn from the accumulated public verdict of practitioners — the recurring praises and the recurring gripes — rather than from any single survey. Two themes dominate. Admirers return to one point above all: Broadest catalog including image, audio and video models. Detractors return to another: Cold starts and its own (non-OpenAI) API shape for LLMs.
We publish neither individual reviews nor reviewer identities; the composite is an editorial synthesis, and the reviewer-note count is an order-of-magnitude indication of how much public signal informs it.
Related dossiers · see also
Peer entries in the same or adjacent category, for comparison against Replicate's standing.