Get ready to compound your AI. August 11.
Save your spotStart free. Scale when you're ready
Try it free. Ship with your team. Scale with ours.
No credit card required
Free
$0
Up to $50 in credits for 3 months
Explore and build your first custom model, no credit card required
- Full automation with Oumi agent (Pat. Pend.)
- Data synthesis with open/closed models
- Data analysis and curation
- Open/closed model evaluation with failure modes
- SFT and PEFT (LoRA, QLoRA) training
For teams and builders
Pro
$25/month
pay-as-you-go after
The power and flexibility to ship real workloads for professional model development
- Everything in Free
- Training with on-policy distillation
- Multiple concurrent jobs for higher efficiency
- Download model weights: 1 / month
- Production inference with autoscaling
- $25/month credits platform-wide
For Production at Scale
Enterprise
Custom
For enterprise scale, with enterprise-grade security & controls – full support from Oumi
- Everything in Pro
- BYOC - VPC & on-prem
- Dedicated capacity and autoscaling to 1000s of GPUs
- Guaranteed throughput and SLAs
- Advanced training methods (RL, custom pipelines)
- Dedicated experts embedded with your team
- Bespoke engagements scoped to your goals
Hosted Platform – detailed pricing
Detailed breakdown of tools, storage, training, and inference pricing.
Tools & Storage
Evaluation | 1,000 judgments / $1 |
Data Synthesis | 1,000 rows / $1 |
Storage | 4 GB/month / $1 |
Supervised Fine-Tuning
Priced per 1M training tokens — calculated as the number of tokens in your training dataset multiplied by the number of epochs.
| Model Size | Price |
|---|---|
Up to 16B | $0.49 |
16.1–32B | $2.00 |
32.1–80B | $3.00 |
80.1–300B | $6.00 |
On-Policy Distillation
Priced per GPU-hour on dedicated GPUs. Training currently runs on 8 GPUs; GPU used is subject to availability.
| GPU | Price / GPU-hr |
|---|---|
A100-80GB | $2.90 |
H100-80GB | $4.00 |
Inference for evaluation & synthesis
| Model Size | Input / 1M | Output / 1M |
|---|---|---|
Llama 3.3 70B | $1.00 | $1.00 |
Qwen2.5 7B Instruct | $0.22 | $0.22 |
Qwen3 235B A22B Instruct 2507 | $0.25 | $0.90 |
Qwen3.5 9B | $0.11 | $0.17 |
Qwen3.5 397B A17B | $0.70 | $4.00 |
Kimi K2.5 | $0.66 | $3.30 |
Kimi K2.6 | $1.05 | $4.40 |
Kimi K3 | $3.30 | $16.50 |
gpt-oss-120b | $0.15 | $0.60 |
DeepSeek-V4-Pro | $1.91 | $3.83 |
GLM-5 | $1.10 | $3.50 |
GLM-5.1 | $1.55 | $4.85 |
GLM-5.2 | $1.55 | $4.85 |
Gemma 4 31B | $0.22 | $0.55 |
Inference is only charged when you utilize models hosted by Oumi to power an action on the platform.
Production Inference
Deploy fully fine-tuned or LoRA models on dedicated GPUs, priced per GPU-hour and billed for uptime. LoRA and full fine-tunes cost the same because each runs on its own dedicated GPU.
| GPU | Price / GPU-hr |
|---|---|
A100-80GB | $3.00 |
H100-80GB | $7.00 |
H200-141GB | $7.00 |
B200-180GB | $10.00 |