Provider Dossier · GPU Cloud
Nebius AI Studio
EU-hosted inference with a full GPU cloud behind it.
The Datum Index verdict
Sound, with caveats
Ranked No. 17 of 20 in our composite reputation index, and currently improving.
≈ 4.0 / 5 aggregate · 170 reviewer notes
Composite index · reviewed 10 July 2026
Findings
- Nebius AI Studio holds a Datum Index composite reputation of 79 / 100 (≈ 4.0 / 5), ranked No. 17 of 20.
- It is a gpu cloud based in Amsterdam, Netherlands, founded 2024, offering open-weight model access.
- Its flagship offering is Open models on EU infrastructure, exposed through an OpenAI-compatible API.
- Indicative pricing is $0.40 / 1M tokens in / $0.60 / 1M tokens out (medium confidence).
- Reference throughput is ~198 tok/s with ~1130 ms time-to-first-token; stated uptime is 99.5%.
Nebius AI Studio enters our index at No. 17, a placement that reads as sound, with caveats. The composite rests on 170 reviewer notes gathered across the public record, and it is best understood not as a score out of ten but as a standing among peers. The trend line is pointed upward: reviewer sentiment has strengthened over recent quarters.
In its favour
EU-hosted inference with a full GPU cloud behind it.
On the numbers that flatter it: open-weight access, 2 named compliance attestations, and a reference throughput of ~198 tokens per second.
Points of concern
Newer entrant still building its reputation.
The caveat a buyer should price in before committing production traffic. As with every entry, we weigh it against the field rather than against perfection.
Incident & Uptime Note
The record of being there
Nebius AI Studio states an availability of 99.5%. Held over a full year, that headline implies on the order of 43.8 hours of cumulative unavailability — a useful sense of scale, though real incidents cluster rather than spread evenly, and a single bad afternoon can outweigh a quiet quarter.
Observatory Panel
Latency & throughput, in distribution
Monthly panels across the models Nebius AI Studio serves medium
GLM 5.2 (FP4)
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~1320 | ~1170 | ~2110 | ~180 | ~191 | ~213 |
| June 2026 | ~1330 | ~1140 | ~2190 | ~191 | ~197 | ~219 |
| July 2026 | ~1260 | ~1180 | ~2290 | ~179 | ~190 | ~208 |
Kimi K3
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~2080 | ~1810 | ~2840 | ~115 | ~123 | ~143 |
| June 2026 | ~2290 | ~1990 | ~3600 | ~106 | ~112 | ~127 |
| July 2026 | ~2080 | ~1840 | ~3210 | ~114 | ~121 | ~130 |
Llama 3.3 70B Base
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~7020 | ~6590 | ~13010 | ~11 | ~12 | ~14 |
| June 2026 | ~7120 | ~6230 | ~10020 | ~12 | ~13 | ~14 |
| July 2026 | ~7210 | ~6370 | ~11260 | ~12 | ~12 | ~13 |
All figures are single-stream, per-request rates as one user would see them — never aggregate accelerator throughput. Serving modes are not comparable with one another: shared serverless APIs batch many tenants per accelerator, dedicated and self-served deployments hand one tenant the whole card, and specialist silicon is a regime of its own. Read each figure within its mode.
No field returns on record for Nebius AI Studio yet — the ledger above opens as soon as the first practitioner submission clears verification.
Key Facts
The record, in brief
| Category | GPU Cloud |
|---|---|
| Headquarters | Amsterdam, Netherlands |
| Founded | 2024 |
| Model access | Open-weight |
| Flagship / reference | Open models on EU infrastructure · Llama 3.3 70B |
| OpenAI-compatible API | Yes |
| Indicative price (in / out) | $0.40 / 1M tokens / $0.60 / 1M tokens medium |
| Reference throughput | ~198 tok/s |
| Reference TTFT | ~1130 ms |
| Stated uptime | 99.5% |
| Compliance | ISO 27001 · GDPR |
| Composite reputation | 79 / 100 · ≈ 4.0 / 5 · 170 notes |
Reviewer Notes
What the field says
The 170 notes behind Nebius AI Studio's standing are drawn from the accumulated public verdict of practitioners — the recurring praises and the recurring gripes — rather than from any single survey. Two themes dominate. Admirers return to one point above all: EU-hosted inference with a full GPU cloud behind it. Detractors return to another: Newer entrant still building its reputation.
We publish neither individual reviews nor reviewer identities; the composite is an editorial synthesis, and the reviewer-note count is an order-of-magnitude indication of how much public signal informs it.
Related dossiers · see also
Peer entries in the same or adjacent category, for comparison against Nebius AI Studio's standing.