-
Sun Yat-sen University
- Guangzhou
-
04:49
(UTC +08:00) - https://gty111.github.io/info/
- https://orcid.org/0009-0005-2979-4486
Highlights
- Pro
Lists (19)
Sort Name ascending (A-Z)
AI
Benchmark
Compiler & DSL
CV & CG
Diffusion
Framework
Hardware
HPC
Instrumention&Reverse&Assemble
LAB
Math
NLP
Operating Systems
Recommendation
ROCM
Simulators
Template & Theme
Tools
Tutorial & Examples
Stars
A framework for efficient model inference with omni-modality models
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…
Lightweight coding agent that runs in your terminal
An asynchronous streaming data management module for efficient post-training.
A framework for few-shot evaluation of language models.
A high-performance and light-weight router for vLLM large scale deployment
slime is an LLM post-training framework for RL Scaling.
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…
Accurate, large-scale, and extensible simulator for LLM inference Systems
Offline optimization of your disaggregated Dynamo graph
Claude Opus 4.6 wrote a dependency-free C compiler in Rust, with backends targeting x86 (64- and 32-bit), ARM, and RISC-V, capable of compiling a booting Linux kernel.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
📰 Must-read papers and blogs on Speculative Decoding ⚡️
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
An Efficient and Versatile Inference Engine for Distributed LLM Serving
FlashMLA: Efficient Multi-head Latent Attention Kernels
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Official PyTorch implementation for "Large Language Diffusion Models"
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
Achieve state of the art inference performance with modern accelerators on Kubernetes
Code for the ICLR 2023 paper "GPTQ: Accurate Post-training Quantization of Generative Pretrained Transformers".
Kimi K2 is the large language model series developed by Moonshot AI team