-
NVIDIA
Stars
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
TokenSpeed is a speed-of-light LLM inference engine.
iTerm2 is a terminal emulator for Mac OS X that does amazing things.
☄🌌️ The minimal, blazing-fast, and infinitely customizable prompt for any shell!
Fast and memory-efficient exact attention
PyTensor allows you to define, optimize, and efficiently evaluate mathematical expressions involving multi-dimensional arrays.
fanshiqing / grouped_gemm
Forked from tgale96/grouped_gemmPyTorch bindings for CUTLASS grouped GEMM.
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Virtual whiteboard for sketching hand-drawn like diagrams
Development repository for the Triton language and compiler
Parse Python docstrings in various flavors.
A Easy-to-understand TensorOp Matmul Tutorial
Remote vanilla PDB (over TCP sockets).
Fast, Flexible and Portable Structured Generation
Enforce the output format (JSON Schema, Regex etc) of a language model
A guidance language for controlling large language models.
A debugging and profiling tool that can trace and visualize python code execution
Demonstration of various hardware effects on CUDA GPUs.
Download M3U8 live streams to the local disk
FlashInfer: Kernel Library for LLM Serving
Universal LLM Deployment Engine with ML Compilation
A high-throughput and memory-efficient inference and serving engine for LLMs