Pricing

A model is a project. A platform is a licence.

Foundry is the work: we build a model and hand you the weights, and the engagement ends there. Studio is a separate product on an annual licence, bought by teams who also want the pipelines around the model. Most customers buy one, not both. Nothing here requires the other.

Bonacci Foundry · the engagement

Model builds

from $15,000 per build

Priced per project by rung, one time. Compute is passed through at cost plus 15% and billed separately, so you can see what the GPUs actually cost. Where the build runs and how much data you bring change the total. The deliverable does not.

Bonacci Studio · optional, separate

Platform licence

from $60,000 a year

Annual, self-hosted, priced on deployment scale and support tier. No quotas, no per-seat AI tax, no metering inside your network. Decline it and a delivered model works exactly the same.

Bonacci Foundry

Three rungs, and picking the right one usually saves you money

Every engagement starts with a one week corpus audit, free for now, that can conclude you should not build a model at all. Roughly half of them do, and you get the written report either way. Across the three rungs, full builds land between $15,000 and $120,000 before compute.

Scoping engagement start here

One week, no charge at present. We audit whatever data you hold, measure tokenizer fit, estimate a token budget, settle where the build should run, and recommend a rung. You get a written report whether or not you continue, and it is yours to take to another vendor.

01 · Fine-tuning

from $15,000 · about 2 weeks. New behaviour, format, or task performance on an open-weights base. Instruction tuning or LoRA, delivered as safetensors and GGUF for your existing vLLM or Ollama setup.

02 · Continued pretraining

from $30,000 · 3 to 5 weeks. New domain knowledge, not just new behaviour. Corpus construction, continued pretraining on an open base, then post-training. The right answer for most teams who arrive asking for rung 3.

03 · Pretraining from scratch

from $60,000 · 6 to 9 weeks. Only correct when your languages tokenize badly on existing vocabularies, your data is not natural text, or every training token must be certifiable. Tokenizer design, corpus pipeline, architecture, pretraining, post-training, eval harness, technical report.

cost + 15%
compute, billed separately
30 / 30 / 40
milestone billing
Yours outright
weights, evals, configs
No
licence to renew on a model

Billing is by milestone: 30% on corpus delivery, 30% at the mid-training checkpoint with evaluation, 40% on final delivery. Compute runs on our GPUs, in your own environment or cloud account, or on isolated hardware inside your air gap, whichever your compliance boundary requires. When it is ours, it is invoiced at what it cost plus 15%. When it is yours, there is nothing for us to bill. We do not mark up GPUs to make the project look cheaper.

What moves the number

Two things set the price, and neither is a tier

The rung sets the floor. These two decide where inside the range you land, and the audit week settles both before anything is committed.

Where the build runs

Inside an air gap: offline install, mirrored images, delivery on physical media, and the extra handling that goes with it. On your hardware: nothing for us to bill on compute, engineering time only. On our GPUs, if nothing requires a boundary: compute at cost plus 15%, billed separately.

How much data you bring

A full corpus is the cheapest path. A few hundred seed examples adds synthetic expansion and the decontamination work that keeps the eval honest. Nothing but the domain adds corpus curation and licence filtering. All three end at the same deliverable.

What does not move it

Seats, users, tokens served, or how well the model performs after you own it. There is no metering on a delivered model because there is nothing of ours left in the path to meter it with.

Book the corpus audit · free See how a build runs
Bonacci Studio

Two ways to run Studio

Studio is bought separately from a model build and neither one requires the other. Self-host it inside your perimeter on an annual licence, or evaluate it in our cloud. Bring your own model either way, including weights we built for you.

Sovereign self-hosted regulated teams

from $60,000 a year. Runs entirely on your infrastructure. Your data, schemas, prompts, and model weights never leave your perimeter. Priced on deployment scale and support tier, invoiced annually.

No quotas on pipelines, connections or executions, because it is your hardware. Local inference against Ollama or vLLM on your own GPUs. Offline licence validated by signature, with no licence server to reach. Your LDAP, Active Directory or OIDC. Your object storage and SMTP. Security review support with your architecture team, a named support engineer, and an agreed response SLA.

Hosted evaluation

Self-serve plans from free. Our infrastructure, for trying it before a deployment call. Appropriate for evaluation and non-sensitive data. If your data cannot leave your network, this is the wrong option and the self-hosted licence is the one to read.

Self-hosted licences are invoiced annually. No payment processor, no metering, and no billing callbacks run inside your network.

Book a deployment call Try the hosted version
Before you ask

What these numbers do not include

Compute

Billed separately at cost plus 15% on a Foundry build. On a Studio deployment you are running your own hardware, so there is nothing for us to bill.

Per-seat AI charges

There are none, on either product. Inference runs against a model server you operate, so usage does not reach us and cannot be metered by us.

A certification we do not hold

We do not hold SOC 2, ISO 27001, or HIPAA attestation. If your procurement requires one today, we are the wrong supplier and would rather say so now. The security page is explicit about this.