Skip to content
View amy-77's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report amy-77

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Memory Sparse Attention - A scalable, end-to-end trainable latent-memory framework for 100M-token contexts.

Python 3,514 226 Updated May 6, 2026

A vLLM plugin built on the FlagOS unified multi-chip backend.

Python 74 99 Updated Aug 9, 2026

[ICLR 2025] MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts

Python 278 12 Updated Oct 16, 2024

A curated list of reinforcement learning, preference optimization, and reward-driven post-training and alignment methods for video generation.

3 Updated Aug 9, 2026

Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers

Jupyter Notebook 44 2 Updated Jul 1, 2026

Self-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.

Python 28,127 2,219 Updated Aug 8, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 88,615 20,487 Updated Aug 10, 2026

The official implementation of NOSA

Python 20 Updated Jun 11, 2026

Train the smallest LM you can that fits in 16MB. Best model wins!

Python 5,177 3,296 Updated May 4, 2026

Official implementation of "WorldKV: Efficient World Memory with World Retrieval and Compression"

Python 89 3 Updated Jul 18, 2026

Code for NeurIPS 2024 paper "Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs"

Python 47 5 Updated Feb 20, 2025

Ongoing research training transformer models at scale

Python 17,380 4,344 Updated Aug 10, 2026

Code for the ICML 2026 Tutorial "Probabilistic Numerics — Computation is Machine Learning"

HTML 57 1 Updated Jul 6, 2026

🔥 [ICML'26] ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs

Python 30 4 Updated Jun 29, 2026

RAT+: Train Dense, Infer Sparse - Recurrence Augmented Attention for Dilated Inference (ICML2026)

Python 10 1 Updated May 21, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 6,911 645 Updated Jul 9, 2026
Python 401 50 Updated Jul 30, 2026

Understand and test language model architectures on synthetic tasks.

Python 282 55 Updated Mar 22, 2026

[ICLR 2024] Efficient Streaming Language Models with Attention Sinks

Python 7,258 399 Updated Jul 11, 2024

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,138 400 Updated Aug 6, 2026

Codes for the paper "∞Bench: Extending Long Context Evaluation Beyond 100K Tokens": https://arxiv.org/abs/2402.13718

Python 388 32 Updated Sep 25, 2024

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

C 21,064 1,891 Updated Aug 9, 2026

Fast CUDA matrix multiplication from scratch

Cuda 1,277 209 Updated Sep 2, 2025

Long Video Gen Infrastructure

Python 2,526 242 Updated Aug 7, 2026

Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).

Python 21 4 Updated Jun 9, 2026

🚀🚀 Efficient implementations of Native Sparse Attention

Python 622 15 Updated Sep 29, 2025

Segmented Code Adjustment Quantization (SAQ)

C++ 27 10 Updated Sep 22, 2025

Query-Adaptive Vector Search

C++ 77 23 Updated Mar 19, 2026

Triton kernels and PyTorch ops for Block Attention Residuals (AttnRes)

Python 87 7 Updated May 29, 2026
Next