Usage changes the cost curve
At enough repeated volume, a dedicated compact model can cost less per completed task than recurring frontier API calls. The comparison has to include real serving utilization.
Nablo takes one production workload from evaluation and approved training data through post-training, proof, and deployment. Your existing product stays in place. You keep the customer-specific model or adapter, subject to base-model and third-party licenses.
Nablo uses existing training and serving infrastructure. We own the workload definition, evaluation, lawful data path, method choice, experiments, proof, and deployment handoff.
The model returns to the same product and workflow. Nablo is not another agent builder, foundation model, or low-level training API.
Frontier APIs are usually the fastest way to prove a product. A bounded production workload becomes a specialization candidate when one of these constraints starts to matter.
At enough repeated volume, a dedicated compact model can cost less per completed task than recurring frontier API calls. The comparison has to include real serving utilization.
A model deployed in your cloud or on dedicated infrastructure gives you more control over response time, capacity, versioning, and availability.
A general model can be capable yet inconsistent on one narrow workflow. Task-specific data and evaluation can improve the behaviour that matters without rebuilding the application.
Training is not automatic and reinforcement learning is not mandatory. The method follows the workload, the evidence, and the evaluation.
We isolate one repeated model-backed task, inspect its tools and failure modes, and freeze the current quality, cost, and latency before changing it.
We use the evidence the workload supports: approved traces, tests, tool results, human corrections, replay, or business outcomes. We fix the prompt or harness first when that is enough.
We choose the appropriate post-training method, compare the result against the current system on the frozen evaluation, and deploy only if the result clears the agreed bar.
Nablo specializes one model-backed component inside the product already running. We do not replace your application, orchestration, tools, or customer experience.
Supervised fine-tuning may be enough. Distillation, preference optimization, or RL are used only when the data and reward support them.
You keep the customer-specific model or adapter and its evaluation. Deploy it in your cloud, or use dedicated Nablo hosting with an agreed export path.
The strongest fit is an AI-enabled B2B software company with a repeated LLM workflow already running in a paid product or contracted service, a technical owner, and a measurable quality, cost, latency, or control problem.
Not a fit: early prototypes, general employee copilots, low-volume workflows, or teams looking for someone to build the entire application.
These are technical experiments, not customer results. They show the evaluation and post-training discipline behind the product.
Qwen3.5-9B matched GPT-5.5 medium's aggregate score on one frozen 220-task benchmark. The post covers the training path, a failed GRPO run, regressions, cost, and limits.
Five experiments across supervised training, soft labels, student-state correction, on-policy probability distillation, and self-distillation. Each post shows the method, result, and failure.
Nablo works directly with production AI teams to establish feasibility, compare a compact model against the system they use today, and deploy only when the result clears the agreed performance bar.
If the workload is not suitable for specialization, we say so early.
Research depth supports the product. The job is to deliver a result that works inside a real system.
Founder and principal engineer
Isaac has a PhD in multi-agent reinforcement learning, more than ten years of enterprise ML experience, and previously co-founded Resoniks, raising €2.65 million.