Skip to content
View GeneZC's full-sized avatar
🌊
Timing
🌊
Timing

Block or report GeneZC

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Minimalistic 4D-parallelism distributed training framework for education purpose

Python 2,274 197 Updated Aug 26, 2025

PyTorch bindings for CUTLASS grouped GEMM.

Cuda 192 50 Updated Apr 8, 2026
Python 615 75 Updated Sep 23, 2025

[ICLR 2026] When it comes to optimizers, it's always better to be safe than sorry

Python 418 14 Updated Sep 26, 2025

Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

Go 178,131 17,325 Updated Aug 9, 2026
Python 92 15 Updated Nov 21, 2025

Survey of Small Language Models from Penn State, ...

259 22 Updated Nov 6, 2025

Chat with multiple PDFs locally

Python 678 107 Updated Oct 23, 2025

A subset of PyTorch's neural network modules, written in Python using OpenAI's Triton.

Python 606 33 Updated May 13, 2026

Quantized Attention on GPU

Python 45 Updated Nov 22, 2024

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,607 464 Updated Jan 17, 2026

The Official Implementation of Ada-KV [NeurIPS 2025]

Python 139 8 Updated Nov 26, 2025

[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Python 540 41 Updated Feb 10, 2025

On-device AI across mobile, embedded and edge for PyTorch

Python 4,879 1,101 Updated Aug 9, 2026

LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via Hybrid Architecture

Python 211 17 Updated Jan 6, 2025

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Jupyter Notebook 19,755 1,831 Updated Jan 30, 2026

大模型进阶面经

156 8 Updated Feb 12, 2026

Python Intelligence Config Manager. A superset of hydra+pydantic+lsp

C 28 Updated Sep 6, 2025

FlagGems is an operator library for large language models implemented in the Triton Language.

Python 1,067 485 Updated Aug 9, 2026

Odysseus: Playground of LLM Sequence Parallelism

Python 83 8 Updated Jun 17, 2024

A framework for serving and evaluating LLM routers - save LLM costs without compromising quality

Python 5,316 415 Updated Aug 10, 2024
Python 4,711 469 Updated Jun 15, 2026

An Open Source Toolkit For LLM Distillation

Python 998 132 Updated May 12, 2026

A family of compressed models obtained via pruning and knowledge distillation

384 21 Updated Nov 6, 2025

BitBLAS is a library to support mixed-precision matrix multiplications, especially for quantized LLM deployment.

Python 770 60 Updated Aug 6, 2025

Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.

Python 2,104 118 Updated Jul 29, 2024
Python 7 Updated Jun 26, 2024

The official evaluation suite and dynamic data release for MixEval.

Python 254 40 Updated Nov 10, 2024

source code of paper "On the Hallucination in Simultaneous Machine Translation"

Python 2 Updated Jun 1, 2024

[Neurips2024] Source code for xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token

Jupyter Notebook 184 20 Updated Jul 4, 2024
Next