Kiminohappydaysinference Get a key

Pricing

Per token. Nothing else.

No seats, no monthly minimum, no egress fees. You pay for tokens processed and audio handled. Spend caps are hard limits — hit one and requests return 429 rather than quietly costing more.

Plans

Start free, scale by usage.

Free

$0 / month

  • 1M tokens per month
  • 20 requests / minute
  • All models except the 235B and 671B flagships
  • Community support

Pay as you go

Usage · from $0.07 / 1M

  • Every model, both regions
  • 600 requests / minute
  • Per-key budgets and spend caps
  • Same-day email support

Reserved

Custom

  • Dedicated GPU capacity
  • Region pinning, EU or US only
  • 99.9% uptime in contract
  • Invoicing in EUR or USD, DPA on request

Language models

Token rates

USD per million tokens. Input covers prompt tokens; output covers generated tokens including reasoning tokens on R1.
Model IDContextInput / 1MOutput / 1M
qwen3-235b-a22b256K$0.70$2.40
deepseek-v3128K$0.50$1.50
deepseek-r1128K$0.55$2.19
mistral-small-3.1128K$0.10$0.30
devstral-small128K$0.10$0.30
gemma-3-27b128K$0.12$0.30
phi-416K$0.07$0.14

Audio & retrieval

Priced by natural unit

Audio and retrieval models are not billed per LLM token.
Model IDTaskUnitPrice
whisper-large-v3Transcriptionper minute of audio$0.004
kokoro-82mSpeech synthesisper 1M characters$0.80
bge-m3Embeddingsper 1M input tokens$0.02
bge-reranker-v2-m3Rerankingper 1M tokens scored$0.05
Billing notes. Usage meters in near real time and is visible per key in the dashboard. Cached prompt prefixes are billed at 10% of the input rate. Batch jobs (24h window) run at half price. Invoices issue monthly; reserved capacity can be prepaid.