lil is a standalone Go launcher for local and Spark/RDMA vLLM deployments in
the local inference lab. It reads launch manifests from the catalog
repository local-inference-lab/lil-catalog, reads each model's checkpoint
metadata from the repository that holds its weights, combines them with a
discovered machine topology, and produces a typed vLLM invocation. Shell
launcher scripts, Ray, Python wrappers, and external queueing layers are not
part of the execution path.
git clone git@github.com:local-inference-lab/lil.git
cd lil
make check
make install
lil --versionThe default install path is ~/.local/bin/lil. Set PREFIX to install
elsewhere:
make install PREFIX=/usr/localThe executable embeds the model families in configs/models/_bases.yaml.
Catalog entries and checkpoint facts come from Hugging Face at runtime, and
topology YAML is generated locally under ~/.config/lil/topologies.
A topology records hardware facts and host tuning. Local discovery records the vLLM checkout, CUDA installation, B12X checkout, GPU memory, compute capability, every valid local GPU pool, and the environment variables a launch on this host exports:
lil discover local \
--repo-root ~/projects/vllm-hh-rebase \
--b12x-root ~/projects/b12xSpark/RDMA discovery runs from the controller and reaches ranks over SSH. The controller does not need to be a Spark node:
lil discover spark \
--node tachyon \
--node luxon \
--repo-root ~/projects/vllm-hh-rebase \
--b12x-root ~/projects/b12xRepeated --node flags define rank order. Discovery validates remote GPU,
RDMA, interface, path, image, and cache compatibility before writing YAML. Use
--output - to inspect the result or --force to replace an existing topology.
Every topology carries two policy fields that discovery fills with defaults:
default_tpselects how a launch picks its tensor-parallel size when--tpis absent.fitchooses the smallest size whose sharded weights leave KV headroom on one rank;alluses every rank. Local discovery writesfit, Spark discovery writesall, and--default-tpoverrides either.environmentholds the host tuning exported by every launch on that topology: allocator settings, thread counts, NCCL protocol and PCIe all-reduce policy on local hosts, loader buffer sizes and RDMA routing on Spark nodes. The block is required, so a topology written before it existed must be rediscovered or edited. Values the launcher derives from topology facts, such asCUDA_HOME,CUTE_DSL_ARCH,CUDA_VISIBLE_DEVICES,LD_LIBRARY_PATH, and the per-rank interface variables, cannot appear in it.
Discovery also records nvrtc_library_dir, the venv directory holding the
CUDA 13 NVRTC builtins, when the runtime Python ships them. Launches that
select the Humming MoE backend put it on LD_LIBRARY_PATH and fail closed
when it is absent.
List the catalog's models, their stored size, and the default tensor-parallel size on every discovered topology:
lil listPretty-print the exact environment and vLLM command without downloading model weights or executing it:
lil render GLM-5.3-NVFP4 --config local --tp 8Run preflight validation. Local checks construct the vLLM engine configuration with the configured interpreter to validate model, scheduler, and parallelism settings before any workers start:
lil check GLM-5.3-NVFP4 --config local --tp 8Update the standard Hugging Face cache with visible progress and launch:
lil run GLM-5.3-NVFP4 --config local --tp 8Local data parallel serving is implemented through vLLM's internal load
balancer, with one API endpoint. --dp always defaults to 1; set it explicitly
to launch multiple replicas:
lil run Qwen3.8-27B-NVFP4-QAD --config local --tp 2 --dp 4This allocates eight GPUs from a discovered device pool. --devices may specify
the full ordered list; each consecutive group of TP devices belongs to one
replica. TP selection and memory estimates remain per replica, and TP * DP
must fit the topology. --max-num-seqs and --max-num-batched-tokens also apply
per replica. lil render --format json reports both parallel sizes and all
selected devices. Data parallel sizes greater than one are unsupported on
Spark/RDMA topologies.
For DP greater than one, the inherited prefill_compute_share and
prefill_compute_half_life defaults are omitted because vLLM does not support
compute sharing across DP ranks. Explicit family or manifest values are
rejected; set both fields to null or use DP=1. Other scheduler settings retain
their configured values.
For Spark/RDMA, lil run updates the cache on every selected rank, repeats the
per-node host probes so that a conflict that appeared during the download is
caught, then starts workers and the head:
lil run GLM-5.3-Flash-NVFP4-Spark \
--config spark \
--tp 2 \
--detach--config local and --config spark select the unique discovered topology of
that kind. A topology name or YAML path selects an exact configuration.
Use an explicit checkpoint directory without changing the model profile:
lil run GLM-5.3-NVFP4 \
--config local \
--model-path /data/models/GLM-5.3-NVFP4 \
--tp 8The local checkpoint must declare the same architectures and attention heads
as the repository at its resolved commit, and its MTP expert quantization must
agree with any assertion in the manifest. Its safetensors index sizes the
memory estimate. For a Spark topology, --sync-model incrementally copies an
explicit --model-path to every rank.
Arguments after -- are forwarded verbatim. Launcher-managed vLLM flags,
including --revision, are rejected there so the command cannot contain
contradictory policy:
lil render GLM-5.3-NVFP4 --tp 8 -- --disable-log-requests--env NAME=VALUE overrides a topology or manifest tuning value for one
launch. Variables the launcher derives, and the offline switches it clears,
cannot be overridden.
Catalog and launch commands resolve the catalog repository's head commit through the
Hugging Face API on each invocation and read each requested manifest from its
local override when present, or from that commit otherwise. They resolve the
entry's weight repository at its pinned revision, or at head when unpinned,
and read config.json and the size of
every safetensors shard at that commit. Bytes already cached for an immutable
commit are reused. Network resolution errors fail the command rather than
silently using stale model policy.
For DFlash, the draft's catalog entry must name the draft repository and list the serving model as compatible.
An unpinned entry never emits --revision, and lil run invokes
hf download OWNER/REPOSITORY without one: the HF CLI checks repository head,
downloads missing or changed files into the standard cache, and streams its
progress to the terminal. A pinned entry emits --revision and downloads that
commit. DFlash launches update both the serving model and the draft
repository. An explicit --model-path skips the serving-model download but
still updates any Hub-backed draft.
For Spark/RDMA, lil updates each rank's cache through SSH. It prefers the hf
executable beside the topology's runtime Python, then checks PATH, with
HF_HOME set to the configured cache mount source. If neither exists, it tries
the configured Docker image's hf executable with that cache mounted. If no
hf executable is available, lil prints a warning and continues because vLLM
can still download during startup; that fallback may not expose useful
progress. A present hf command returning an authentication, network, or
filesystem error fails the run.
lil import registers an existing model directory under its Hugging Face
repository ID without downloading file contents or duplicating weight storage:
lil import local-inference-lab/Qwen3.8-Flash-Next-NVFP4 \
/data/models/qwen3.8-flash-next-mixed/qwen3.8-flash-next-180b-nvfp4-ple-mxfp8-attn-shared_vv1 \
--dry-runRemove --dry-run to write the cache entries. Pass the directory containing
the repository files, such as config.json and weight shards. This command
accepts a repository ID directly and does not require a catalog entry or
topology. --revision SHA selects a 40-character commit; otherwise it resolves
main and updates the cache's refs/main after importing.
The importer reads Hub metadata and verifies every matching local file's size
and content hash before writing cache entries: SHA-256 for LFS weights and
Git blob SHA-1 for regular files. It hardlinks files into blobs/ and creates
relative symlinks in snapshots/<commit>/. Existing cache files must match;
conflicts fail the import. Unrelated local files and incomplete downloads are
left alone. Missing repository files are listed explicitly; a later hf download or lil run can fetch them while reusing the imported files.
The source and cache must share a filesystem. Hardlinks keep the data alive if either path is deleted, but editing a linked file in place changes both paths: treat imported files as immutable. There is no copy fallback.
--cache-dir selects the Hub cache directory. Its default follows
HF_HUB_CACHE, HUGGINGFACE_HUB_CACHE, then HF_HOME/hub, with
XDG_CACHE_HOME/huggingface/hub or ~/.cache/huggingface/hub as the fallback.
LIL_CACHE_DIR controls catalog metadata and does not change the Hub cache.
Verification progress and counts go to stderr; stdout contains the snapshot
path. A dry run performs the same content checks without writing cache entries.
EDITOR=vim lil edit GLM-5.3-NVFP4
EDITOR='code --wait' lil edit GLM-5.3-NVFP4 --repositorylil edit NAME opens the existing local override, or copies
NAME/lil.yaml from local-inference-lab/lil-catalog when no override exists.
Saving a change writes ~/.config/lil/models/NAME/lil.yaml (under
$XDG_CONFIG_HOME when set). This complete manifest shadows the upstream
entry for Hub-backed list, render, check, run, and cluster commands. Family
defaults still apply. An explicit --models-config directory takes precedence
and does not use these overrides. Remove the local file to use upstream again.
Editing an existing override requires no network; Hub-backed launches still
resolve the catalog and checkpoint metadata online.
--repository always opens the upstream manifest and uploads changed contents
to the catalog's main branch on successful editor exit. It uses HF_TOKEN,
HUGGING_FACE_HUB_TOKEN, or the token saved by hf auth login, including
HF_TOKEN_PATH, HF_HOME, and XDG_CACHE_HOME token locations. The token must
have write access to local-inference-lab/lil-catalog. Uploads use a parent
commit precondition to reject concurrent repository changes. Local overrides
remain in place and still shadow an uploaded manifest; the command reports
when one exists.
EDITOR is required and may contain quoted arguments. GUI editors must use
their wait option. Manifests are validated before saving or uploading; unchanged
files produce no write. Editor, validation, and upload failures retain the
working YAML file and report its path for recovery. Draft catalog entries can
also be edited.
Every launchable model is one entry in local-inference-lab/lil-catalog: a
directory named for the model holding a lil.yaml. The directory name is the
launch name. The manifest names the repository that holds the weights and
states only what the checkpoint cannot state about itself. Architectures,
attention heads, stored weight size, and the quantization of the MTP expert
layer are read from config.json and the shard sizes at the resolved commit.
schema_version: 1
kind: model
model: local-inference-lab/GLM-5.3-NVFP4
family: glm
description: GLM-5.3 with NVFP4 routed experts and an unquantized BF16 MTP expert layer
serving:
served_model_name: GLM-5.3
trust_remote_code: true
async_scheduling: true
generation_config: vllm
long_prefill_token_threshold: 2048
speculators:
default: mtp
mtp:
tokens: 3
moe_quantization: bf16
attention: B12X
draft_sample_method: probabilistic
model: target
compilation:
cudagraph_mode: FULL_AND_PIECEWISE
custom_ops:
- all
environment:
CUDA_DEVICE_MAX_CONNECTIONS: "32"Precision, cache policy, backend selection, and loading are separate mappings:
precision:
dtype: bfloat16
quantization: modelopt_mixed
cache:
kv_cache_dtype: fp8
indexer_kv_dtype: auto
block_size: 16
mamba_cache_mode: align
recurrent_checkpoint_policy: request_boundaries
prefix_cache_retention_interval: 4096
kernels:
attention: B12X
linear: b12x
moe: b12x
gdn_prefill: b12x
gdn_decode: b12x
kda_prefill: auto
mamba: triton
flashinfer_autotune: true
loading:
load_format: b12x
loader_extra_config: {}Set only the options relevant to the model and hardware. String selectors must be non-empty; vLLM validates backend and dtype support at runtime. Unknown section keys are rejected. Each mapping inherits family values; explicit manifest values win, and null clears inheritance for that value. Defaulted fields then use launcher defaults, while optional fields emit no argument.
The sections:
modelnames the weight repository in owner/name form; it may belong to any account.revisionoptionally pins a commit. A repository outsidelocal-inference-labthat enablestrust_remote_codemust be pinned, because a third-party head is not a trusted input. A pinned entry emits--revision, downloads that commit, reads its checkpoint facts there, and carries the pin into a speculative config whose draft lives in the target checkpoint.familynames one entry in the embedded families file. Family mappings merge under the manifest; a scalar, list, or explicitnullin the manifest replaces the family value.servingholds the API-facing policy: served name, remote code, tokenizer mode, parsers, tool choice, generation config, default chat-template arguments, scheduler switches including the prefill schedule interval, usage and request-ID reporting, and multimodal limits.auto_tool_choicedefaults to true when a tool parser is set. Prefix caching and chunked prefill default to on. A manifest without parsers or a tool parser emits none of those flags.precisionholdsdtype(default:bfloat16) and optionalquantization. These select compute precision and the checkpoint quantization implementation.cacheholdskv_cache_dtype(default:fp8),indexer_kv_dtype,block_size,mamba_cache_mode,recurrent_checkpoint_policy, andprefix_cache_retention_interval.--kv-cache-dtypeoverrides the manifest for one launch.indexer_kv_dtypeacceptsauto,bf16,fp8,mxfp4, ornvfp4in the configured vLLM fork and is passed in--attention-config.recurrent_checkpoint_policyacceptsauto,aligned, orrequest_boundaries. Request boundaries retain the verified leading instruction prefix, completed prompt, and committed response endpoint; aligned retention uses block boundaries.autoselects request boundaries when supported and aligned retention otherwise. The retention interval is a non-negative integer;0retains only semantic checkpoints. Omission leaves the vLLM default in effect.block_sizemust be a positive integer.kernelsselectsattention,linear,moe,gdn_prefill,gdn_decode,kda_prefill, andmambabackends and controlsflashinfer_autotune. Linear and MoE backends default tob12x. In the configured vLLM fork,gdn_prefillacceptsflashinfer,triton,cutedsl, orb12x;gdn_decodeacceptsb12x,cuda, ortriton;kda_prefillacceptsauto,triton,flashkda, orb12x; andmambaacceptstriton,flashinfer, orcpu. B12X GDN prefill and decode must be selected consistently. Optional backend selectors emit no option when omitted or null.loadingholdsload_formatand theloader_extra_configmapping. Local PCIe launches default toinstanttensor; Spark (SM121) launches default tob12xwith theb12x_loaderplugin enabled.loading.load_formatoverrides that platform default. Extra configuration is passed as JSON through--model-loader-extra-config.speculatorsnames adefaultofmtp,dflash,dspark, ornone(the default) and a section per method.mtp.moe_quantizationis an assertion checked against the checkpoint; it is required only when the checkpoint metadata cannot decide, andmtp.moe_backendoverrides the backend the quantization implies. NVFP4, MXFP4, and BF16 experts use the B12X backend, MXFP8 experts use Humming, and block-FP8 experts need an explicit backend. A Humming launch adds the topology's discovered NVRTC library directory, which the Humming kernels load at runtime. A DSpark draft always ships inside the target checkpoint.dspark.adaptive_verificationsets its default verification policy;--adaptive-verification=trueor--adaptive-verification=falseoverrides it for one launch.capacity,compilation, andenvironmentare the base launch layer. Capacity defaults aremax_model_len: auto, eight sequences, and 4096 batched tokens;capacity.gpu_memory_utilizationreplaces the topology's default for this model. Every capacity value is a default that the matching command-line flag overrides. A compilation mapping enables--compilation-configwith computed CUDA graph capture sizes.overridesis an ordered list of layers applied when every condition inwhenmatches the launch. Conditions arekind(localorspark_rdma),arch(the topology's CuTe DSL architecture, such assm_121a),tp, andspeculator(the method that actually runs, with a zero-depth launch counting asnone). Later overrides win.requires.archlists the architectures a checkpoint may run on. A launch on any other topology fails, andlil listmarks the topology.
Draft checkpoints are entries of kind: draft. They are not listed as
serving models; a serving entry's speculators.dflash.model must match a
draft entry whose compatibility list contains the serving model's repository:
schema_version: 1
kind: draft
model: local-inference-lab/GLM-5.3-Flash-DFlash2-MXFP8
description: MXFP8 DFlash speculative-decoding draft for GLM-5.3 Flash
method: dflash
quantization: mxfp8
compatible_models:
- local-inference-lab/GLM-5.3-Flash-NVFP4--models-config DIRECTORY reads a local catalog laid out the same way, with
each model entry's config.json and model.safetensors.index.json beside its
manifest, for development and tests. The fixtures under
internal/launcher/testdata/model-manifests are the reference copies of the
published catalog. Unknown keys, missing families, inheritance cycles, and
invalid types fail closed.
| Concern | Owner |
|---|---|
| Architectures, attention heads, stored size, MTP expert quantization | weight repository config.json and shard sizes at the resolved commit |
| Weight repository, pinned revision, serving identity, parsers, precision, cache policy, kernel selection, loading, speculation, capacity and compilation layers, conditional overrides | catalog entry plus its family |
| GPU inventory, CUDA path, memory-utilization ceiling, default TP policy, host tuning environment, API bind defaults, local pools, SSH/RDMA layout, image, and cache mounts | discovered topology YAML |
| TP, local checkpoint, bind overrides, KV dtype, scheduler limits, capacity, profiling, speculation overrides, and one-off environment overrides | typed lil CLI |
| Derived environment, default TP fit, graph capture sizes, native commands | Go launcher |
| Loaded checkpoint truth, runtime imports, GPU state, paths, ports, HCAs, and vLLM argument compatibility | preflight checks |
A launch may use any tensor-parallel size the topology can host that divides
the checkpoint's attention heads. Local launches take the first device pool
large enough; Spark launches take the first N one-GPU ranks.
The topology owns the default --gpu-memory-utilization, with 0.95 written
by discovery; a catalog entry may replace it and the command line overrides
both.
CUDA graph capture sizes are calculated from resolved scheduler and speculation
settings. The set contains mixed batch sizes and every uniform decode batch
through max_num_seqs; with MTP or DSpark depth K, uniform verification
shapes are sequence count multiplied by K+1. The mixed ladder extends to twice
the uniform maximum, except for DSpark, whose verifier never exceeds one
sampled token plus its drafts per request.
The stored weight size lets lil estimate post-sharding device weights per
rank, mapped-host PLE bytes for an explicit checkpoint, a runtime reserve, and
an advisory safe KV budget before downloading a checkpoint. The same
arithmetic picks the default tensor-parallel size under the fit policy. JSON
rendering exposes every term, the supported TP sizes, the checkpoint facts and
their source, and the manifest commit:
lil render GLM-5.3-NVFP4 --tp 8 --format jsonBy default, lil emits the topology's --gpu-memory-utilization and leaves
--kv-cache-memory-bytes absent so vLLM profiles the actual model, activation
peak, and graph footprint. --kv-cache-memory-bytes requests a fixed allocation;
auto restores runtime profiling.
Rank containers are named deterministically from the topology prefix, model,
and TP, and they are kept after exit. A crashed rank leaves its logs and exit
code behind; the next launch removes an exited container of the same name and
refuses a running one. Preflight parses each rank's arguments inside the launch
image with the launch mounts and environment, so the interpreter and imports
validated are the ones the container runs. Spark preflight also hashes the
importable vllm/ and b12x/ Python trees on the controller and every
selected rank; native extensions are not compared.
Without --detach, the controller starts workers, then the head, streams the
head's logs, and polls every rank. It tears the cluster down only when it
observes a rank exit or is interrupted. A worker exit stops the head and prints
the worker's last log lines; a head exit returns its exit code. Losing contact
with the ranks never stops a running cluster: the controller reports the
outage, keeps trying for ten minutes, then gives up and leaves the containers
running for lil cluster to manage. Ctrl-C stops and removes every rank; a
second Ctrl-C exits immediately.
With --detach, the controller confirms every rank is still running a few
seconds after start and returns.
lil cluster status GLM-5.3-Flash-NVFP4-Spark --config spark --tp 2
lil cluster logs GLM-5.3-Flash-NVFP4-Spark --config spark --tp 2 --follow
lil cluster logs GLM-5.3-Flash-NVFP4-Spark --config spark --tp 2 --rank 1
lil cluster wait GLM-5.3-Flash-NVFP4-Spark --config spark --tp 2 --timeout 30m
lil cluster profile-start GLM-5.3-Flash-NVFP4-Spark --config spark --tp 2
lil cluster profile-stop GLM-5.3-Flash-NVFP4-Spark --config spark --tp 2
lil cluster stop GLM-5.3-Flash-NVFP4-Spark --config spark --tp 2cluster stop stops and removes every rank. The controller performs lifecycle
and profiler requests over SSH with keepalives. The head API does not need to
be exposed to the controller network.
lil run --sync-code reconciles the two remote Python package directories
with rsync --delete, then reruns preflight; without that flag, lil never
changes remote source trees.