-
WestlakeU
- Hangzhou, Zhejiang, PRC
- https://quancs.github.io
Stars
The agent that grows with you
OpenClaw-RL: Train any agent simply by talking
A PyTorch native platform for training generative AI models
Ongoing research training transformer models at scale
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance…
A safetensors extension to efficiently store sparse quantized tensors on disk
Official PyTorch implementation for "Large Language Diffusion Models"
[FSE'2026] SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
[ICLR 2025] COAT: Compressing Optimizer States and Activation for Memory-Efficient FP8 Training
PyTorch native quantization for training and inference
A collection of research papers on low-precision training methods
FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.
Efficient Triton Kernels for LLM Training
Bridge Megatron-Core to Hugging Face/Reinforcement Learning
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Community maintained hardware plugin for vLLM on Ascend
A high-throughput and memory-efficient inference and serving engine for LLMs
A real browser preview inside your editor that you can debug.
TinyNeuralNetwork is an efficient and easy-to-use deep learning model compression framework.
TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation
Pytorch implementation of "CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR".
Development repository for the Triton language and compiler
Fast and memory-efficient exact attention
🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
MoBA: Mixture of Block Attention for Long-Context LLMs
Code for NeurIPS 2024 paper - The GAN is dead; long live the GAN! A Modern Baseline GAN - by Huang et al.
Several simple examples for popular neural network toolkits calling custom CUDA operators.
Homogram is a 3rd-party Telegram client for HarmonyOS 5, driven by ArkTS/ArkUI (UI-layer) and Rust (native-layer).