Skip to content
View bingps's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report bingps

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Lean 165 22 Updated Aug 17, 2026

DeepGEMM: clean and efficient BLAS kernel library on GPU

Cuda 7,692 1,180 Updated Aug 11, 2026

Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict Mechanism (NIPS'25)

Python 16 Updated Oct 6, 2025

Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

Cuda 11,774 1,240 Updated Aug 17, 2026

[NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive

Cuda 75 6 Updated Aug 8, 2026

🚀🚀 Efficient implementations of Native Sparse Attention

Python 603 15 Updated Sep 29, 2025

🚀 Efficient implementations for emerging model architectures

Python 5,569 662 Updated Aug 17, 2026

Nano vLLM

Python 15,035 2,469 Updated Apr 26, 2026

诺亚盘古大模型研发背后的真正的心酸与黑暗的故事。

11,542 1,300 Updated Jul 9, 2025

iTerm2 + Oh My Zsh 打造舒适终端体验

1,876 302 Updated Feb 4, 2026

CPM.cu is a lightweight, high-performance CUDA implementation for LLMs, optimized for end-device inference and featuring cutting-edge techniques in sparse architecture, speculative sampling and qua…

Cuda 242 27 Updated Jan 14, 2026

SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Python 216 20 Updated Jul 10, 2026

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,649 482 Updated Jan 17, 2026

Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).

Python 2,501 297 Updated Feb 20, 2026

🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"

Python 1,019 54 Updated Feb 5, 2026

FlashInfer: Kernel Library for LLM Serving

Python 6,181 1,298 Updated Aug 18, 2026

📰 Must-read papers on KV Cache Compression (constantly updating 🤗).

733 28 Updated Aug 16, 2026

A throughput-oriented high-performance serving framework for LLMs

Jupyter Notebook 974 52 Updated Mar 29, 2026

Recipes to scale inference-time compute of open models

Python 1,132 131 Updated May 26, 2026

[MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention

C++ 853 68 Updated Mar 6, 2025

[SIGMOD 2025] PQCache: Product Quantization-based KVCache for Long Context LLM Inference

Python 91 25 Updated Dec 7, 2025

LLM KV cache compression made easy

Python 1,172 168 Updated Aug 17, 2026

Free, open source crypto trading bot

Python 53,395 11,103 Updated Aug 18, 2026

A curated list of insanely awesome libraries, packages and resources for Quants (Quantitative Finance)

HTML 28,926 3,866 Updated Aug 18, 2026

[ICLR2025 Spotlight] MagicPIG: LSH Sampling for Efficient LLM Generation

Python 256 21 Updated Dec 16, 2024

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filli…

Python 1,226 81 Updated Apr 8, 2026

Tile primitives for speedy kernels

Cuda 3,631 319 Updated Jul 13, 2026

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6,303 1,090 Updated Aug 17, 2026

ArkVale: Efficient Generative LLM Inference with Recallable Key-Value Eviction (NIPS'24)

Python 54 9 Updated Dec 17, 2024

Source code of "FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework"

Cuda 11 3 Updated Oct 23, 2024
Next