Enpar Inference
Inference, served in Asia-Pacific.
Open models, or your own, served on GPUs in India, Malaysia, Indonesia and Singapore. Billed per million tokens, or per GPU-hour on capacity of your own.
What you get
Close to your users. Sized to your traffic.
In the region
Four countries to serve fromYour requests run on GPUs in India, Malaysia, Indonesia or Singapore, under the data-protection law of the country they run in.
Capacity
Room to growOur own GPUs first. When your traffic needs more, we add capacity we have checked, on the same contract and the same invoice.
Models
Open models, or yoursOpen-weight models served for you, or your own fine-tuned weights on GPUs reserved for you.
Three ways to run
Pick how you want to serve.
- Per token
Per million tokens, in and out
Call an open model and pay for what you send and what comes back. Nothing to set up and nothing to reserve.
- Dedicated
Per GPU-hour
Your model, or a model you choose, on GPUs that serve nobody else's traffic. For fine-tuned weights and steady load.
- Reserved
Per GPU-hour, for a term
Capacity held for you for a fixed term, priced by the length of the term. For launches and traffic you can forecast.
Where we run
Four countries. Each under its own law.
Our GPUs run in India, Malaysia, Indonesia and Singapore, and each country’s data-protection law applies where its GPUs are.
- India
- Digital Personal Data Protection Act 2023
- Malaysia
- Personal Data Protection Act 2010, amended 2024
- Indonesia
- Personal Data Protection Law No. 27 of 2022
- Singapore
- Personal Data Protection Act 2012
When you need more
Your traffic grows. Your contract doesn’t change.
We serve from our own GPUs first. When demand runs past them, we place the overflow on capacity from data centres we have checked, and it stays on your one contract and your one monthly invoice from Enpar.
- First
- Our own GPUs
- Then
- Capacity we have checked
- Always
- One contract, one invoice
Models
The models we serve.
The catalogue is published here as models go live. Tell us the model you need, and we confirm in writing whether we serve it, where, and at what price.
Questions
What buyers ask first.
- Where do our requests run?
- On GPUs in India, Malaysia, Indonesia or Singapore, under the data-protection law of the country they run in.
- Which models can we run?
- Open-weight models we serve, or your own weights on dedicated GPUs. Tell us the model, and we confirm in writing.
- How is it billed?
- Per million tokens in and out, or per GPU-hour for dedicated and reserved capacity. One invoice a month, from Enpar.
- What if we need more capacity than you have?
- We add capacity from data centres we have checked. Your contract and your invoice stay the same.
- When are prices public?
- Soon, on our homepage, beside the market rate. Until then we quote in writing, and we reply within one working day.