-
diNo Research Group, LIPADE Lab
- Beijing, China
-
10:53
(UTC +08:00) - https://amy-77.github.io/
- in/yanlin-qi-456177268
Stars
Memory Sparse Attention - A scalable, end-to-end trainable latent-memory framework for 100M-token contexts.
A vLLM plugin built on the FlagOS unified multi-chip backend.
[ICLR 2025] MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
A curated list of reinforcement learning, preference optimization, and reward-driven post-training and alignment methods for video generation.
Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers
Self-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.
A high-throughput and memory-efficient inference and serving engine for LLMs
Train the smallest LM you can that fits in 16MB. Best model wins!
Official implementation of "WorldKV: Efficient World Memory with World Retrieval and Compression"
Code for NeurIPS 2024 paper "Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs"
Ongoing research training transformer models at scale
Code for the ICML 2026 Tutorial "Probabilistic Numerics — Computation is Machine Learning"
🔥 [ICML'26] ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
RAT+: Train Dense, Infer Sparse - Recurrence Augmented Attention for Dilated Inference (ICML2026)
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
Understand and test language model architectures on synthetic tasks.
[ICLR 2024] Efficient Streaming Language Models with Attention Sinks
SkyRL: A Modular Full-stack RL Library for LLMs
Codes for the paper "∞Bench: Extending Long Context Evaluation Beyond 100K Tokens": https://arxiv.org/abs/2402.13718
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
Fast CUDA matrix multiplication from scratch
Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).
🚀🚀 Efficient implementations of Native Sparse Attention
Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes)