Milestones
List view
Guide an operator from a fresh management checkout and supplied private inventory to Narwhal running on their GPU hardware. The public documentation covers management access, remote installation, engine and fabric preparation, process-bound attestation, fleet configuration, fresh profiling, preflight, a routed request, measured workload and working Prometheus/Grafana observability. Each step names its host, required inputs, command, expected result and recovery action. Site addresses, credentials, launch records and raw deployment evidence stay in private files. Fleet size comes from the supplied engine inventory. The milestone finishes when the documented private-route trial passes, client outcomes reconcile with the router journal, monitoring shows the router and engines, and the post-load KV ring passes. #3 tracks final documentation consistency and integration.
No due date•11/12 issues closedAn operator supplies a model revision, accelerator type and GPU budget, representative input/output lengths, target request rate, and TTFT/TPOT budgets. Narwhal enumerates supported homogeneous-TP engine shapes and prefill/decode splits, profiles operator-provided engines where measurements are missing, and projects capacity for each candidate from the measured curves. Each measurement carries the model revision, runtime image, hardware, launch settings, probe domain, and raw samples. Startup and P/D transfer checks determine candidate eligibility; the search places workloads beyond the measured domain in a separate unqualified section. For a target load, the report ranks shapes by GPUs needed to meet the latency budgets and shows throughput headroom, source measurements, and a draft fleet configuration. Operators apply the draft through their existing deployment workflow. End-to-end load at two request rates compares projected and observed TTFT, TPOT, and throughput for the recommended shape. The report records prediction error and the workload range supported by the measurements. Completion requires a reproducible comparison of at least two fleet shapes for the same workload. Mixed-TP shapes enter the search after pair compatibility and shape-aware scheduling are qualified.
No due dateAn operator selects `hardware.tensor_parallel` and `hardware.accelerators_per_engine` for a homogeneous fleet, then launches each engine with the matching TP size. Narwhal compares those declared values with process-bound launch evidence, profiles every engine, and gates serving on preflight and P/D transfer checks. The first qualification pairs TP1 with a model that fits one GPU. A multi-GPU shape follows the same procedure so the guide covers an actual choice of TP sizes. For each shape, the deployment guide names the fleet fields and launcher settings, explains recovery from launch mismatches and failed profiles, and records the measured workload range. A two-rate load run records TTFT, TPOT, and throughput for each qualified shape. A shape qualifies when launch evidence matches the fleet config, profiles cover the intended workload, preflight accepts the fleet, P/D handoff succeeds, and end-to-end load meets the stated SLOs. Every engine in a qualifying fleet uses the same TP size; mixed-TP pairing and scheduling will be addressed in future work.
No due date