Pricing

Every price, published.

What you pay, before you talk to anyone. US dollars, before tax; we can also bill in [Insert: local currencies].

01 · Model API

Pay per token.

Per million tokens, by model. [Insert: launch models — six candidates shown]

ModelForInput · 1MOutput · 1M
gpt-oss-120bReasoningUS$[Insert]US$[Insert]
DeepSeek V4 FlashChatUS$[Insert]US$[Insert]
GLM-5.3-FlashChatUS$[Insert]US$[Insert]
Llama 3.1 8BSmall and fastUS$[Insert]US$[Insert]
SEA-LIONSoutheast Asian languagesUS$[Insert]US$[Insert]
BGE-M3EmbeddingsUS$[Insert]—

Prices in US dollars, before tax

All models

02 · Dedicated endpoint

Pay per GPU-hour.

Your model, or one of ours, on GPUs reserved for you.

GPUPer hour
H100 80 GBUS$[Insert]
H200 141 GBUS$[Insert]

[Insert: the GPUs we run — two candidates shown]

03 · Reserved capacity

GPUs held for you, under contract.

Private and in-country if your data needs it. By quote.

Hardware
Single-tenant GPUs in [Insert: city]
For
Steady volume, private deployments, regulated data
Scale
More than one hall when you need it

Not sure which?

Tell us what you want to run and we will tell you what it costs.

Talk to our team