Skip to content
View WangXuan95's full-sized avatar

Block or report WangXuan95

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

ModelEngine 项目群的社区管理规范。

Python 50 Updated Jul 9, 2026

[VLDB 26, NeurIPS 25] Scalable long-context LLM decoding that leverages sparsity—by treating the KV cache as a vector storage system.

Python 152 36 Updated Sep 14, 2026

High performance block-sorting data compression library

C 352 63 Updated Feb 3, 2026

High-speed lossless data compression of 16 to 512 bytes--get better average compression than QuickLZ for 512-byte blocks. td512 maintains good compression down to 16-byte blocks.

C 27 Updated Feb 14, 2022

ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization

Python 114 19 Updated Oct 15, 2024

Daily updated LLM papers. 每日更新 LLM 相关的论文,欢迎订阅 👏 喜欢的话动动你的小手 🌟 一个

Python 1,330 60 Updated Sep 16, 2026

Awesome-LLM-KV-Cache: A curated list of 📙Awesome LLM KV Cache Papers with Codes.

466 29 Updated Jun 17, 2026
Python 333 35 Updated Jul 10, 2025

MQSim is a fast & accurate simulator for modern multi-queue (MQ) and SATA SSDs. MQSim faithfully models new high-bandwidth protocol implementations, steady-state SSD conditions, and full end-to-end…

C++ 390 198 Updated Aug 25, 2025
C++ 77 16 Updated May 30, 2023

Open Source SSD Controller. NVMe and Lightstor variants

Bluespec 18 23 Updated May 21, 2014

现代图形引擎入门指南

C++ 482 55 Updated Dec 16, 2025

Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Cuda 247 24 Updated Sep 24, 2023

Opensource DDR3 Controller

Verilog 481 80 Updated Sep 12, 2026

Using LLM to evaluate MMLU dataset.

Python 42 3 Updated Mar 8, 2024

A collection of benchmarks and datasets for evaluating LLM.

581 38 Updated Jul 13, 2024

qoi and qoi-like implementations optionally using simd

C 14 1 Updated Nov 28, 2024

The simplest, fastest repository for training/finetuning medium-sized GPTs.

Python 63,323 10,886 Updated Nov 12, 2025

📰 Must-read papers on KV Cache Compression (constantly updating 🤗).

744 29 Updated Aug 16, 2026

AISystem 主要是指AI系统,包括AI芯片、AI编译器、AI推理和训练框架等AI全栈底层技术

Jupyter Notebook 17,882 2,482 Updated Sep 3, 2025

KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches. EMNLP Findings 2024

Python 90 3 Updated Feb 27, 2025

Fast LZMA2 Library

C 346 34 Updated Sep 12, 2026

翻译systemverilog assertion部分

8 1 Updated Jul 15, 2024

Insane(ly slow but wicked good) PNG image optimization

Python 3,424 149 Updated Jun 18, 2022

cmix is a lossless data compression program aimed at optimizing compression ratio at the cost of high CPU/memory usage.

C++ 728 57 Updated Sep 2, 2026

Trabalho de Graduação

C++ 20 4 Updated Nov 2, 2014

A random event driven text-based game engine.

TypeScript 264 39 Updated Aug 19, 2024

state-of-the-art lossless audio compression

C++ 64 6 Updated Sep 22, 2026
Next