Research lab · Self-deployed frontier models

Frontier models. Your hardware.

Make self-deploying open-source models cheaper, easier and faster. We choose, adapt and deploy the model around your workload and the cards you can afford.

For product teams using paid model APIs in production, operators running agents around the clock, and companies whose data must stay on premises. For teams already running open models or preparing to deploy them on their own hardware or cloud account.

Choose the work you need

Start with the decision or build in front of you.

See all services

Research you can inspect

The lab researches engines, quantization, pruning, speculative decoding and Hebrew models. Published experiments describe their own setup and limits; they are not promises about your deployment.

Browse the research

Shared scale from zero; units: %

Code replay

Code replay: BF16 62.61 %; NVFP4 61.32 %BF1662.61 %NVFP461.32 %
First-draft agreement on code replay. (%)Measured on our reference setup on Hardware: RTX PRO 6000 BlackwellQwen3.5-9B draft agreement; not live throughput.Receipt hqmtp
Inspect the data table
First-draft agreement on code replay. · %
ConditionBF16NVFP4
Code replay62.61 %61.32 %

Compression is what lets a model fit a smaller card. In our code replay, compressed and uncompressed runs showed similar first-draft agreement, a speed-technique measurement, not answer quality. The measurement record has the setup and results.

What fits your hardware?

A model fitting in memory is only the start. The choice also depends on answer quality, request size, simultaneous work and the cost of operating the system.

Task and examplesQuality checksExpected loadHardware and budgetDecision factorsModel choiceMemory and executionResponse timeTotal operating costOutputA measuredrecommendation and adeployment scope
  1. Inputs

    • Task and examples
    • Quality checks
    • Expected load
    • Hardware and budget
  2. Decision factors

    • Model choice
    • Memory and execution
    • Response time
    • Total operating cost
  3. Output

    • A measured recommendation and a deployment scope.

Discuss an assessment

Does the move pay?

We compare your API bill with the full cost of running a suitable open model on your hardware or cloud account. The proposal covers adaptation, deployment, infrastructure, operation and support, with setup and running costs shown separately. It shows whether the move saves money. You control the deployment, data and logs.

From examples to a system you can run

Each stage has a clear purpose, from the first comparison to handover and continuing support.

  1. Choose

    Define the task and compare model and hardware options.

  2. Adapt

    Test fine-tuning, architecture and engine changes where they address a measured gap.

  3. Deploy

    Install, test and document the system on your hardware or cloud account.

  4. Support

    Support coverage and handover are set out in your proposal.

Our engineers keep supporting the systems they build, under agreed terms. About the lab →

Bring the workload and the budget.

We will work out what to measure and what to build.