Talk to an engineer

Sovereign inference.Run any model.

Run any of 300+ open AI models on infrastructure you control, at up to 70% lower cost.*

Xinference
Admin
Documentation
AAdmin
v2.3.0
Overview
Cluster resources and model deployment overview · Last updated 14:40:51
Launch a model
Add a model to your cluster. Existing endpoints keep running.
Quick pick
Gemma 4Chat · LLM
Qwen3.5 0.8BChat · LLM
Dedicated Models
3
Hosted Models
3
Requests
2.4K
API requests · this billing period
Running Bill
$2,155.12
USD · this billing period
Request Volume · Past 24h
15:0021:0003:0009:00Now
Model Utilization
See all →
Dedicated Models
GLM 562.0K req/hr
Llama 3.3 70B21.0K req/hr
DeepSeek V433.5K req/hr
Hosted Models
qwen2-72b45.2K req/hr
llama3.1-70b32.1K req/hr

Interactive preview for illustration. Sign in to the console for the real product.

Trusted in production.

Siemens
Yum!
AIA
Everbright Securities
TFC OpticalComms
Berry Genomics
XW Bank
0%

Runs in your environment. Your data never leaves.

UPTO0%*

Lower cost than closed-source APIs.

00×

Faster responses, more models per GPU, no cold starts.

*Savings estimate reflects self-hosted or on-premises deployment of open models compared with closed-source API list pricing. Actual savings vary by workload and utilisation.

One control plane, from app to GPU.

One integration. It keeps working as models, hardware and traffic change.

Your apps

One OpenAI-compatible API. No migration.

Xinference

Routes every request to the right model and GPU, automatically.

Smart routingGPU poolingAutoscalingObservability

Models & GPUs

300+ open models, plus your own fine‑tuned models, on any hardware.

One line to switch over
base_url = "http://your-xinference:9997/v1"

Run any of 300+ open models, from day one.

Llama 3.1Qwen 3.6gpt-oss-120bDeepSeek V4MistralKimi K2GemmaGLM 5bge-large-en+300 more

Our cloud, your cloud, or on‑prem.

One standard bundle to get started, or a custom deployment in your own environment.

Standard bundle · Fully managed

Everything to run sovereign AI: $10K per month.

One plan with a dedicated private LLM, hosted in Australia, ready to go.

  • Unlimited AI agents
  • 500 users included
  • Dedicated private LLM, Australian hosted
See pricing →
Custom deployments · Your environment

On your own infrastructure, cloud or on-premises.

Tailored deployments for teams that need Xinference inside their own perimeter.

  • No data crosses your perimeter
  • Tailored to your requirements
  • Best for regulated data: finance, health, government
Talk to an engineer →

Where the savings come from.

No trade-off on quality, just hardware working harder.

Squeeze more from the hardware.

KV-cache reuse, continuous batching, speculative decoding keep GPUs busy.

Up to 4× more from every GPU

Run open models.

Same quality bar, a fraction of the price.

Up to 70% lower cost*
Powered by XRouter

Route intelligently.

Each request goes to the cheapest model that clears the quality bar.

Cheapest capable model, always

Your data never leaves.

No third-party transit
Private by design
Isolated compute
SSO · SAML
Audit logs
Governed onboarding
Data residency AU · SG
Role-based access
API key controls
Live monitoring (TTFT · TPOT)
End-to-end encryption
Financial servicesHealthcareGovernmentInsuranceIndustrial operations

Keep everything you already have.

Pilot it next to what you run today. Expand once the numbers hold up.

01

Proof of concept

1–2 weeks

Start on the standard bundle. Nothing to build.

02

Benchmark

With your baseline

Cost, accuracy and latency, measured against what you run today. We run it with you.

03

Production

Once proven

Move to private cloud or on-prem once the numbers hold up.

Get started

Run any model.
With Xinference.

Sovereign by default, lower cost by design, faster in production.