Let's help help help devs.
Target: <1000ms. Several important LLM CLI tools take multiple seconds. PRs welcome ❤️
First upstream <1s attempt: vLLM PR #41518.
Run: Jul 16, 2026 12:00 UTC / CPU: Core™ i7-8700K; GPU: 1x RTX 5060 Ti
tensorrt-llm --help
vllm --help
VLMEvalKit --help
sglang --help
transformers --help
datasets --help
llm --help
openai --help
hf --help
lm-eval --help
langchain-cli --help
tokenspeed --help
ollama --help
llama.cpp --help