Kiminohappydaysinference Get a key

Catalogue

Models

Every model here is open-weight and served on our own hardware — nothing is proxied to a third-party API. Model IDs are the strings you pass in the model field. Prices are per million tokens unless noted; see pricing for the full table and rate limits.

Language

Chat & reasoning

Nine language models. Context is the maximum window we serve; some cards advertise more than we host stably.
Model IDMakerBest forContextLicenceInOut
qwen3-235b-a22b Alibaba General, reasoning, tools 256K Apache 2.0 $0.70$2.40
deepseek-v3 DeepSeek General workhorse 128K MIT $0.50$1.50
deepseek-r1 DeepSeek Reasoning, maths, hard code 128K MIT $0.55$2.19
mistral-small-3.1 Mistral AI Fast multimodal (vision) 128K Apache 2.0 $0.10$0.30
devstral-small Mistral AI Coding agents, SWE tasks 128K Apache 2.0 $0.10$0.30
gemma-3-27b Google Efficient multimodal 128K Gemma Terms $0.12$0.30
phi-4 Microsoft Small, strong reasoning 16K MIT $0.07$0.14

Audio

Speech to text, text to speech

Both audio models are billed on their natural unit rather than tokens: transcription per minute of input, synthesis per million characters of output.

whisper-large-v3

MIT

OpenAI · speech-to-text

Transcription and X→English translation across 99 languages. Returns segments or word-level timestamps; handles noisy, accented, and code-switched audio well.

Endpoint/audio/transcriptionsLanguages99 Timestampsword / segmentPrice$0.004 / min

kokoro-82m

Apache 2.0

Hexgrad · text-to-speech

Lightweight, natural 24 kHz synthesis with 54 voices across 8 languages. Low latency to first audio; streams PCM or returns WAV/MP3/Opus.

Endpoint/audio/speechVoices54 Sample rate24 kHzPrice$0.80 / 1M ch

Retrieval

Embeddings & reranking

For RAG pipelines. Embeddings are priced per million input tokens; reranking per million tokens scored.
Model IDMakerTaskDim / contextLicencePrice
bge-m3 BAAI Embeddings (multilingual) 1024 / 8K MIT $0.02 / 1M
bge-reranker-v2-m3 BAAI Cross-encoder rerank — / 8K Apache 2.0 $0.05 / 1M
Model IDs are stable. When we retire a snapshot we keep the ID pointed at the newest compatible weights and give 30 days’ notice by email and on the status page. Pin a dated snapshot (e.g. deepseek-v3-0324) if you need bit-for-bit reproducibility.