Provider Dossier · Subscription (flat-rate)
Cheapestinference
Flat-rate subscription pricing that decouples cost from token volume, in continuous operation since September 2025.
The Datum Index verdict
Recommended with confidence
Ranked No. 11 of 20 in our composite reputation index, and currently improving.
≈ 4.2 / 5 aggregate · 96 reviewer notes
Composite index · reviewed 10 July 2026
Findings
- Cheapestinference holds a Datum Index composite reputation of 83 / 100 (≈ 4.2 / 5), ranked No. 11 of 20.
- It is a subscription (flat-rate) based in United States, founded 2025, offering open-weight model access.
- Its flagship offering is Flat-rate subscription tiers (Core / Frontier / Flagship), exposed through an OpenAI-compatible API.
- Pricing is a flat monthly subscription — reserved 8-hour daily blocks with unlimited tokens, tiers from $14.99 to $149 per block per month — with no per-token metering (medium confidence).
- Reference throughput is ~88 tok/s with ~1400 ms time-to-first-token; stated uptime is 99.5%.
Cheapestinference enters our index at No. 11, a placement that reads as recommended with confidence. The composite rests on 96 reviewer notes gathered across the public record, and it is best understood not as a score out of ten but as a standing among peers. The trend line is pointed upward: reviewer sentiment has strengthened over recent quarters.
In its favour
Flat-rate subscription pricing that decouples cost from token volume, in continuous operation since September 2025.
On the numbers that flatter it: open-weight access, 1 named compliance attestation, and a reference throughput of ~88 tokens per second.
Points of concern
Youngest entrant in the index — operating since September 2025 with a clean record to date; the only thing left to accumulate is tenure.
The caveat a buyer should price in before committing production traffic. As with every entry, we weigh it against the field rather than against perfection.
Incident & Uptime Note
The record of being there
Cheapestinference states an availability of 99.5%. Held over a full year, that headline implies on the order of 43.8 hours of cumulative unavailability — a useful sense of scale, though real incidents cluster rather than spread evenly, and a single bad afternoon can outweigh a quiet quarter.
Observatory Panel
Latency & throughput, in distribution
Monthly panels across the models Cheapestinference serves medium
Kimi K2.7
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~1580 | ~1410 | ~2510 | ~82 | ~87 | ~97 |
| June 2026 | ~1570 | ~1470 | ~2660 | ~81 | ~84 | ~96 |
| July 2026 | ~1590 | ~1500 | ~2910 | ~77 | ~82 | ~90 |
GLM 5.2
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~1430 | ~1260 | ~2110 | ~96 | ~101 | ~115 |
| June 2026 | ~1390 | ~1290 | ~2070 | ~92 | ~99 | ~113 |
| July 2026 | ~1460 | ~1260 | ~2480 | ~98 | ~101 | ~110 |
DeepSeek V4 Flash
| Window | TTFT ms — mean · p50 · p90 | Throughput tok/s — mean · p50 · p90 | ||||
|---|---|---|---|---|---|---|
| May 2026 | ~1220 | ~1140 | ~2400 | ~107 | ~110 | ~118 |
| June 2026 | ~1340 | ~1250 | ~2120 | ~98 | ~101 | ~115 |
| July 2026 | ~1450 | ~1260 | ~2030 | ~93 | ~100 | ~113 |
All figures are single-stream, per-request rates as one user would see them — never aggregate accelerator throughput. Serving modes are not comparable with one another: shared serverless APIs batch many tenants per accelerator, dedicated and self-served deployments hand one tenant the whole card, and specialist silicon is a regime of its own. Read each figure within its mode.
No field returns on record for Cheapestinference yet — the ledger above opens as soon as the first practitioner submission clears verification.
Key Facts
The record, in brief
| Category | Subscription (flat-rate) |
|---|---|
| Headquarters | United States |
| Founded | 2025 |
| Model access | Open-weight |
| Flagship / reference | Flat-rate subscription tiers (Core / Frontier / Flagship) · Kimi K3 |
| OpenAI-compatible API | Yes |
| Pricing model | Pricing is a flat monthly subscription — reserved 8-hour daily blocks with unlimited tokens, tiers from $14.99 to $149 per block per month — with no per-token metering medium |
| Reference throughput | ~88 tok/s |
| Reference TTFT | ~1400 ms |
| Stated uptime | 99.5% |
| Compliance | GDPR |
| Composite reputation | 83 / 100 · ≈ 4.2 / 5 · 96 notes |
Reviewer Notes
What the field says
The 96 notes behind Cheapestinference's standing are drawn from the accumulated public verdict of practitioners — the recurring praises and the recurring gripes — rather than from any single survey. Two themes dominate. Admirers return to one point above all: Flat-rate subscription pricing that decouples cost from token volume, in continuous operation since September 2025. Detractors return to another: Youngest entrant in the index — operating since September 2025 with a clean record to date; the only thing left to accumulate is tenure.
We publish neither individual reviews nor reviewer identities; the composite is an editorial synthesis, and the reviewer-note count is an order-of-magnitude indication of how much public signal informs it.
Related dossiers · see also
Peer entries in the same or adjacent category, for comparison against Cheapestinference's standing.