Kiminohappydaysinference Get a key

Open weights,
served straight.

Kiminohappydays runs the leading open-weight language and audio models on dedicated hardware in the EU and US. One OpenAI-compatible endpoint, billed per token, with nothing trained on your traffic.

No card to start · Frankfurt + Ashburn regions · SOC 2 in progress

client SDK / curl edge TLS · auth · quota router model select batcher continuous gpu pool H100 · H200 stream out SSE tokens request path · p50 to first token ≈ 210 ms
All systems operational p50 TTFT 210 ms p99 640 ms 30-day uptime 99.97% models live 11

Why Kiminohappydays

The open ecosystem, without the ops.

Open-weight models caught up. Running them well — batching, quantisation, tail latency, region residency — is still a job. We do that part, and keep the interface identical to what you already use.

Compatible

Drop-in schema

Chat, embeddings, reranking and audio all follow the OpenAI request shape. Change the base URL and key; keep your SDK.

Private

Nothing retained

Prompts and outputs are never logged beyond the request or used for training. EU traffic stays in the EU when you pin the region.

Honest cost

Per token, hard caps

No seats, no minimums. Spend caps fail closed with a 429 instead of quietly running up a bill.

Selected models

Frontier open weights, plus audio.

A curated set rather than a dump — the models we can serve at low tail latency and stand behind in production. Full catalogue with specs and licences on the models page.

Qwen3-235B-A22B

Apache 2.0

Alibaba · mixture-of-experts

Flagship general and reasoning model. 235B total parameters, 22B active per token, strong multilingual and tool use.

Context256K$ / 1M in$0.70

DeepSeek-R1

MIT

DeepSeek · reasoning

Explicit chain-of-thought reasoning for maths, logic and hard code. Emits thinking tokens you can show or hide.

Context128K$ / 1M in$0.55

Whisper large-v3

MIT

OpenAI · speech-to-text

Robust transcription and translation across 99 languages, with word-level timestamps. Billed by the minute of audio.

Languages99$ / min$0.004

See all 11 models →

First request

Point your client at us.

If your code already speaks the OpenAI schema, this is the whole migration.


        

Read the full documentation →