Skip to content
View AniZpZ's full-sized avatar

Block or report AniZpZ

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Frequently updated list of dLLM (Diffusion Large Language Models) papers, models, and other resources

Python 53 6 Updated Jul 19, 2026
Python 392 26 Updated Apr 16, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 6,930 646 Updated Jul 9, 2026

A compiler, optimizer and executor for financial expressions and factors

C++ 316 60 Updated May 29, 2026

TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration

Python 1,720 192 Updated Mar 27, 2026

A kernel library written in tilelang

Python 1,717 154 Updated Apr 23, 2026

An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Python 566 137 Updated Aug 12, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,058 109,192 Updated Aug 6, 2026

PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.

Python 222 41 Updated Dec 24, 2025

IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse

132 11 Updated Mar 14, 2026

A lightweight inference engine supporting speculative speculative decoding (SSD).

Python 986 78 Updated May 10, 2026

SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models

Python 781 324 Updated Aug 12, 2026

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,132 1,586 Updated Aug 12, 2026

Autonomous GPU Kernel Generation & Optimization via Deep Agents

Python 508 85 Updated Jul 15, 2026

CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

Python 1,126 95 Updated Jul 8, 2026
Python 187 34 Updated Aug 11, 2026

DFlash: Block Diffusion for Flash Speculative Decoding

Python 5,604 403 Updated May 10, 2026

Nsight Python is a Python kernel profiling interface based on NVIDIA Nsight Tools

Python 286 21 Updated Aug 12, 2026

Our first fully AI generated deep learning system

Python 635 49 Updated Feb 2, 2026

High Performance LLM Inference Operator Library

C++ 1,106 131 Updated Aug 6, 2026

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Python 4,587 350 Updated Jan 14, 2026

Open-source unified multimodal model

Python 6,145 546 Updated May 4, 2026

[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Python 540 41 Updated Feb 10, 2025

Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding

Python 117 9 Updated Dec 2, 2025

Accelerating MoE with IO and Tile-aware Optimizations

Python 739 95 Updated Jul 4, 2026

SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention

Python 329 20 Updated Feb 24, 2026

Tile-Based Runtime for Ultra-Low-Latency LLM Inference

Python 1,669 119 Updated Aug 6, 2026
Next