Pricing
Every price, published.
What you pay, before you talk to anyone. US dollars, before tax; we can also bill in [Insert: local currencies].
01 · Model API
Pay per token.
Per million tokens, by model. [Insert: launch models — six candidates shown]
| Model | For | Input · 1M | Output · 1M |
|---|---|---|---|
| gpt-oss-120b | Reasoning | US$[Insert] | US$[Insert] |
| DeepSeek V4 Flash | Chat | US$[Insert] | US$[Insert] |
| GLM-5.3-Flash | Chat | US$[Insert] | US$[Insert] |
| Llama 3.1 8B | Small and fast | US$[Insert] | US$[Insert] |
| SEA-LION | Southeast Asian languages | US$[Insert] | US$[Insert] |
| BGE-M3 | Embeddings | US$[Insert] | — |
Prices in US dollars, before tax
All models02 · Dedicated endpoint
Pay per GPU-hour.
Your model, or one of ours, on GPUs reserved for you.
| GPU | Per hour |
|---|---|
| H100 80 GB | US$[Insert] |
| H200 141 GB | US$[Insert] |
[Insert: the GPUs we run — two candidates shown]
03 · Reserved capacity
GPUs held for you, under contract.
Private and in-country if your data needs it. By quote.
- Hardware
- Single-tenant GPUs in [Insert: city]
- For
- Steady volume, private deployments, regulated data
- Scale
- More than one hall when you need it
Not sure which?
Tell us what you want to run and we will tell you what it costs.