Highlights
- Pro
Lists (1)
Sort Name ascending (A-Z)
Stars
TokenSpeed is a speed-of-light LLM inference engine.
High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.
Quickly setup the infrastructure to run a collaborative autoresearch project
Curated collection of papers in MoE model inference
DeepEP: an efficient expert-parallel communication library
LLM algorithm practice lab with theory, solutions, and test cases.《大模型算法与系统教程》面向大模型入门到进阶的算法实战教程,覆盖原理讲解、答案解析、测试用例与 CUDA/Triton 实战。
ncnn is a high-performance neural network inference framework optimized for the mobile platform
CUDA Python: Performance meets Productivity
LEAKED SYSTEM PROMPTS FOR CHATGPT, CLAUDE, GEMINI, GROK, PERPLEXITY, CURSOR, LOVABLE, REPLIT, AND MORE! - AI SYSTEMS TRANSPARENCY FOR ALL! 👐
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
A Python-level JIT compiler designed to make unmodified PyTorch programs faster.
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…
One local control plane for every AI agent: route across models, fuse new capabilities, orchestrate tools, and stay fully in control.
Input: a vLLM fork url, Output: a vLLM docker image
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Give your agents the power of the Hugging Face ecosystem
[ICLR 2024] Efficient Streaming Language Models with Attention Sinks
Best DDoS Attack Script Python3, (Cyber / DDos) Attack With 56 Methods