Skip to content
View xwqtju's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report xwqtju

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

FlashKDA: high-performance Kimi Delta Attention kernels

Cuda 1,222 118 Updated Jul 30, 2026

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,659 484 Updated Jan 17, 2026

TurboDiffusion: 100–200× Acceleration for Video Diffusion Models

Python 3,615 276 Updated Aug 5, 2026

AI-agent Skill for generating polished HTML slide decks: editorial magazine and Swiss layouts, image prompts, social covers, and a WebGL/low-power presentation runtime.

HTML 127 45 Updated Jul 14, 2026
AGS Script 1 Updated Jun 10, 2026

DeepGEMM: clean and efficient BLAS kernel library on GPU

Cuda 7,706 1,187 Updated Aug 11, 2026

FlashInfer: Kernel Library for LLM Serving

Python 6,197 1,308 Updated Aug 20, 2026

Claude Code 中文全面上手指南。基于 luongnv89/claude-howto 本土化重写,面向中国小白用户,保留命令与配置兼容性,并附学习路径与本地化校验护栏。

Python 2,281 330 Updated Aug 6, 2026

DeepSeek-V4 Lecture

Python 27 4 Updated Aug 10, 2026

分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等

Jupyter Notebook 3,596 346 Updated Aug 7, 2026

Ascend PyTorch adapter (torch_npu). Mirror of https://gitcode.com/Ascend/pytorch

Python 565 84 Updated Aug 20, 2026

Fast and memory-efficient exact attention

Python 24,745 3,000 Updated Aug 19, 2026

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

Python 14,425 2,678 Updated Aug 20, 2026

Repo for Qwen Image Finetune

Jupyter Notebook 1 Updated Dec 12, 2025

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,311 1,618 Updated Aug 20, 2026

Train transformer language models with reinforcement learning.

Python 19,110 2,918 Updated Aug 20, 2026

🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.

Python 21,563 2,439 Updated Aug 19, 2026

Enjoy the magic of Diffusion models!

Python 12,973 1,272 Updated Aug 20, 2026

Repo for Qwen Image Finetune

Jupyter Notebook 253 27 Updated Aug 11, 2026

[NeurIPS 24 Spotlight] MaskLLM: Learnable Semi-structured Sparsity for Large Language Models

Python 189 17 Updated Jan 1, 2025

Officiel code for PATCH: Learnable Tile-level Hybrid Sparsity for LLMs

Python 8 Updated Jul 24, 2026

[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.

Cuda 1,030 102 Updated Feb 25, 2026

🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )

Python 4,798 459 Updated Jul 27, 2026

[ICML 2025] Official PyTorch implementation of "FlatQuant: Flatness Matters for LLM Quantization"

Python 227 35 Updated Nov 25, 2025

Fast and memory-efficient exact attention

Python 37 2 Updated Dec 2, 2024

LLM Finetuning with peft

Jupyter Notebook 2,978 771 Updated Aug 1, 2025

Awesome list for LLM quantization

Python 439 29 Updated Apr 20, 2026

Awesome LLM compression research papers and tools.

1,862 130 Updated Jun 30, 2026
Next