Enpar Inference

Inference, served in Asia-Pacific.

Open models, or your own, served on GPUs in India, Malaysia, Indonesia and Singapore. Billed per million tokens, or per GPU-hour on capacity of your own.

A-06 · CUSTOMER A · BILLED

What you get

Close to your users. Sized to your traffic.

In the region

Four countries to serve from

Your requests run on GPUs in India, Malaysia, Indonesia or Singapore, under the data-protection law of the country they run in.

Capacity

Room to grow

Our own GPUs first. When your traffic needs more, we add capacity we have checked, on the same contract and the same invoice.

Models

Open models, or yours

Open-weight models served for you, or your own fine-tuned weights on GPUs reserved for you.

Three ways to run

Pick how you want to serve.

  1. Per token

    Per million tokens, in and out

    Call an open model and pay for what you send and what comes back. Nothing to set up and nothing to reserve.

  2. Dedicated

    Per GPU-hour

    Your model, or a model you choose, on GPUs that serve nobody else's traffic. For fine-tuned weights and steady load.

  3. Reserved

    Per GPU-hour, for a term

    Capacity held for you for a fixed term, priced by the length of the term. For launches and traffic you can forecast.

Where we run

Four countries. Each under its own law.

Our GPUs run in India, Malaysia, Indonesia and Singapore, and each country’s data-protection law applies where its GPUs are.

India
Digital Personal Data Protection Act 2023
Malaysia
Personal Data Protection Act 2010, amended 2024
Indonesia
Personal Data Protection Law No. 27 of 2022
Singapore
Personal Data Protection Act 2012
INDIAINRMALAYSIAMYRSINGAPORESGDINDONESIAIDR

When you need more

Your traffic grows. Your contract doesn’t change.

We serve from our own GPUs first. When demand runs past them, we place the overflow on capacity from data centres we have checked, and it stays on your one contract and your one monthly invoice from Enpar.

First
Our own GPUs
Then
Capacity we have checked
Always
One contract, one invoice

Models

The models we serve.

The catalogue is published here as models go live. Tell us the model you need, and we confirm in writing whether we serve it, where, and at what price.

Questions

What buyers ask first.

Where do our requests run?
On GPUs in India, Malaysia, Indonesia or Singapore, under the data-protection law of the country they run in.
Which models can we run?
Open-weight models we serve, or your own weights on dedicated GPUs. Tell us the model, and we confirm in writing.
How is it billed?
Per million tokens in and out, or per GPU-hour for dedicated and reserved capacity. One invoice a month, from Enpar.
What if we need more capacity than you have?
We add capacity from data centres we have checked. Your contract and your invoice stay the same.
When are prices public?
Soon, on our homepage, beside the market rate. Until then we quote in writing, and we reply within one working day.

Tell us what you run. A reply in one working day.