Pricing
Per token. Nothing else.
No seats, no monthly minimum, no egress fees. You pay for tokens processed and audio handled. Spend caps are hard limits — hit one and requests return 429 rather than quietly costing more.
Plans
Start free, scale by usage.
Free
$0 / month
- 1M tokens per month
- 20 requests / minute
- All models except the 235B and 671B flagships
- Community support
Pay as you go
Usage · from $0.07 / 1M
- Every model, both regions
- 600 requests / minute
- Per-key budgets and spend caps
- Same-day email support
Reserved
Custom
- Dedicated GPU capacity
- Region pinning, EU or US only
- 99.9% uptime in contract
- Invoicing in EUR or USD, DPA on request
Language models
Token rates
| Model ID | Context | Input / 1M | Output / 1M |
|---|---|---|---|
| qwen3-235b-a22b | 256K | $0.70 | $2.40 |
| deepseek-v3 | 128K | $0.50 | $1.50 |
| deepseek-r1 | 128K | $0.55 | $2.19 |
| mistral-small-3.1 | 128K | $0.10 | $0.30 |
| devstral-small | 128K | $0.10 | $0.30 |
| gemma-3-27b | 128K | $0.12 | $0.30 |
| phi-4 | 16K | $0.07 | $0.14 |
Audio & retrieval
Priced by natural unit
| Model ID | Task | Unit | Price |
|---|---|---|---|
| whisper-large-v3 | Transcription | per minute of audio | $0.004 |
| kokoro-82m | Speech synthesis | per 1M characters | $0.80 |
| bge-m3 | Embeddings | per 1M input tokens | $0.02 |
| bge-reranker-v2-m3 | Reranking | per 1M tokens scored | $0.05 |
Billing notes. Usage meters in near real time and is visible per key in the dashboard. Cached prompt prefixes are billed at 10% of the input rate. Batch jobs (24h window) run at half price. Invoices issue monthly; reserved capacity can be prepaid.