Foundry is the work: we build a model and hand you the weights, and the engagement ends there. Studio is a separate product on an annual licence, bought by teams who also want the pipelines around the model. Most customers buy one, not both. Nothing here requires the other.
Priced per project by rung, one time. Compute is passed through at cost plus 15% and billed separately, so you can see what the GPUs actually cost. Where the build runs and how much data you bring change the total. The deliverable does not.
Annual, self-hosted, priced on deployment scale and support tier. No quotas, no per-seat AI tax, no metering inside your network. Decline it and a delivered model works exactly the same.
Every engagement starts with a one week corpus audit, free for now, that can conclude you should not build a model at all. Roughly half of them do, and you get the written report either way. Across the three rungs, full builds land between $15,000 and $120,000 before compute.
One week, no charge at present. We audit whatever data you hold, measure tokenizer fit, estimate a token budget, settle where the build should run, and recommend a rung. You get a written report whether or not you continue, and it is yours to take to another vendor.
from $15,000 · about 2 weeks. New behaviour, format, or task performance on an open-weights base. Instruction tuning or LoRA, delivered as safetensors and GGUF for your existing vLLM or Ollama setup.
from $30,000 · 3 to 5 weeks. New domain knowledge, not just new behaviour. Corpus construction, continued pretraining on an open base, then post-training. The right answer for most teams who arrive asking for rung 3.
from $60,000 · 6 to 9 weeks. Only correct when your languages tokenize badly on existing vocabularies, your data is not natural text, or every training token must be certifiable. Tokenizer design, corpus pipeline, architecture, pretraining, post-training, eval harness, technical report.
Billing is by milestone: 30% on corpus delivery, 30% at the mid-training checkpoint with evaluation, 40% on final delivery. Compute runs on our GPUs, in your own environment or cloud account, or on isolated hardware inside your air gap, whichever your compliance boundary requires. When it is ours, it is invoiced at what it cost plus 15%. When it is yours, there is nothing for us to bill. We do not mark up GPUs to make the project look cheaper.
The rung sets the floor. These two decide where inside the range you land, and the audit week settles both before anything is committed.
Inside an air gap: offline install, mirrored images, delivery on physical media, and the extra handling that goes with it. On your hardware: nothing for us to bill on compute, engineering time only. On our GPUs, if nothing requires a boundary: compute at cost plus 15%, billed separately.
A full corpus is the cheapest path. A few hundred seed examples adds synthetic expansion and the decontamination work that keeps the eval honest. Nothing but the domain adds corpus curation and licence filtering. All three end at the same deliverable.
Seats, users, tokens served, or how well the model performs after you own it. There is no metering on a delivered model because there is nothing of ours left in the path to meter it with.
Studio is bought separately from a model build and neither one requires the other. Self-host it inside your perimeter on an annual licence, or evaluate it in our cloud. Bring your own model either way, including weights we built for you.
from $60,000 a year. Runs entirely on your infrastructure. Your data, schemas, prompts, and model weights never leave your perimeter. Priced on deployment scale and support tier, invoiced annually.
No quotas on pipelines, connections or executions, because it is your hardware. Local inference against Ollama or vLLM on your own GPUs. Offline licence validated by signature, with no licence server to reach. Your LDAP, Active Directory or OIDC. Your object storage and SMTP. Security review support with your architecture team, a named support engineer, and an agreed response SLA.
Self-serve plans from free. Our infrastructure, for trying it before a deployment call. Appropriate for evaluation and non-sensitive data. If your data cannot leave your network, this is the wrong option and the self-hosted licence is the one to read.
Self-hosted licences are invoiced annually. No payment processor, no metering, and no billing callbacks run inside your network.
Billed separately at cost plus 15% on a Foundry build. On a Studio deployment you are running your own hardware, so there is nothing for us to bill.
There are none, on either product. Inference runs against a model server you operate, so usage does not reach us and cannot be metered by us.
We do not hold SOC 2, ISO 27001, or HIPAA attestation. If your procurement requires one today, we are the wrong supplier and would rather say so now. The security page is explicit about this.