Skip to content
View llsj14's full-sized avatar

Block or report llsj14

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A unified inference runtime for VLA models.

C++ 157 28 Updated Aug 15, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,905 242 Updated Aug 16, 2026

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…

Rust 461 137 Updated Aug 15, 2026

AlpaSim is an open-source autonomous vehicle simulation platform designed for development and testing of end-to-end AV policies

Python 1,178 148 Updated Aug 12, 2026

A safetensors extension to efficiently store sparse quantized tensors on disk

Python 313 112 Updated Aug 15, 2026

Flash Vision-Language-Action Inference for Autonomous Driving

Python 60 7 Updated Jul 6, 2026

CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies

Rust 76,245 4,795 Updated Aug 15, 2026

Open Source AI Platform - AI Chat with advanced features that works with every LLM

Python 31,610 4,350 Updated Aug 15, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 11,162 1,713 Updated Aug 16, 2026

DSPy: The framework for programming—not prompting—language models

Python 37,243 3,222 Updated Aug 15, 2026

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…

Python 14,730 1,296 Updated Aug 15, 2026

Common recipes to run vLLM

JavaScript 973 375 Updated Aug 16, 2026

Offline optimization of your disaggregated Dynamo graph

Python 405 151 Updated Aug 15, 2026

A Datacenter Scale Distributed Inference Serving Framework

Rust 7,773 1,443 Updated Aug 16, 2026

Distributed MoE in a Single Kernel [NeurIPS '25]

Cuda 281 40 Updated May 5, 2026

High-Performance KV Cache Storage Engine on CXL Shared Memory for LLM Inference

Python 56 4 Updated Aug 13, 2026
Python 319 64 Updated Aug 10, 2026

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6,283 1,085 Updated Aug 15, 2026

A collection of prompts, system prompts and LLM instructions

HTML 5,282 721 Updated Aug 14, 2026

Extracted system prompts from Anthropic - Claude Fable 5, Opus 5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-5.6-Sol, Codex. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Curso…

JavaScript 62,967 10,337 Updated Aug 15, 2026

Causal depthwise conv1d in CUDA, with a PyTorch interface

Python 936 204 Updated Aug 14, 2026

🚀 Efficient implementations for emerging model architectures

Python 5,561 659 Updated Aug 14, 2026

Material for gpu-mode lectures

Jupyter Notebook 6,442 642 Updated Jun 15, 2026

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,917 1,002 Updated Aug 13, 2026

TPU inference for vLLM, with unified JAX and PyTorch support.

Python 407 286 Updated Aug 16, 2026

Easy, Fast, and Scalable Multimodal AI

Python 130 12 Updated Jun 2, 2026

[NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)

Python 225 13 Updated Feb 11, 2026

A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do

1,310 164 Updated Apr 27, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,992 1,380 Updated Aug 5, 2026
Next