-
Notifications
You must be signed in to change notification settings - Fork 0
Home
Athrael Soju edited this page Sep 21, 2026
·
1 revision
Narwhal profiles the fleet's engines, admits each request against measured SLO budgets, assigns prefill and decode to separate engines, and shifts the live split as demand changes while model weights stay resident.
Start with the local path, then follow the pages in order for a real deployment.
| Page | Use it for |
|---|---|
| 1. Get started | Run the complete request path on CPU stubs. |
| 2. Core concepts | Understand engines, request placement, reactive control, and state. |
| 3. Deploy | Connect and verify a GPU fleet. |
| 4. Operate Narwhal | Run ingress, telemetry, failover, and engine maintenance. |
| 5. Troubleshoot a fleet | Respond to overload, engine faults, and router faults. |
| 6. Measure a fleet | Calibrate SLOs and validate the deployment under load. |
| 7. Configuration | Look up fleet fields and defaults. |
| 8. CLI reference | Look up commands, options, and exit behaviour. |
| 9. API and data reference | Integrate HTTP routes, state, metrics, and journals. |
| 10. Observability | Generate targets, verify Prometheus scrapes, and inspect the Grafana dashboard. |