Skip to content
Athrael Soju edited this page Sep 21, 2026 · 1 revision

Narwhal logo and wordmark

Narwhal profiles the fleet's engines, admits each request against measured SLO budgets, assigns prefill and decode to separate engines, and shifts the live split as demand changes while model weights stay resident.

Start with the local path, then follow the pages in order for a real deployment.

Page Use it for
1. Get started Run the complete request path on CPU stubs.
2. Core concepts Understand engines, request placement, reactive control, and state.
3. Deploy Connect and verify a GPU fleet.
4. Operate Narwhal Run ingress, telemetry, failover, and engine maintenance.
5. Troubleshoot a fleet Respond to overload, engine faults, and router faults.
6. Measure a fleet Calibrate SLOs and validate the deployment under load.
7. Configuration Look up fleet fields and defaults.
8. CLI reference Look up commands, options, and exit behaviour.
9. API and data reference Integrate HTTP routes, state, metrics, and journals.
10. Observability Generate targets, verify Prometheus scrapes, and inspect the Grafana dashboard.

Clone this wiki locally