Skip to content
View jinjungyu's full-sized avatar

Block or report jinjungyu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Teams-first Multi-agent orchestration for Claude Code

TypeScript 38,524 3,463 Updated Aug 12, 2026

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…

Python 141,184 22,681 Updated Aug 11, 2026

[EMNLP 2025] AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models

Python 16 2 Updated Apr 29, 2026

Code for MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization

Python 23 3 Updated Feb 18, 2026

[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

Python 739 77 Updated May 14, 2026

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Python 5,449 431 Updated Jul 26, 2026

PyTorch native quantization for training and inference

Python 2,941 589 Updated Aug 12, 2026

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

Python 1,225 199 Updated Aug 11, 2026

Open Machine Learning Compiler Framework

Python 13,664 3,952 Updated Aug 11, 2026

Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.

Python 6,246 573 Updated Aug 22, 2025

An implementation of local windowed attention for language modeling

Python 503 49 Updated Jul 16, 2025

Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)

Python 403 48 Updated Apr 22, 2025

Go ahead and axolotl questions

Python 12,344 1,401 Updated Aug 12, 2026

CoreNet: A library for training deep neural networks

Jupyter Notebook 7,005 540 Updated Oct 9, 2025

Explorations into some recent techniques surrounding speculative decoding

Python 307 24 Updated Dec 22, 2024

[ICLR 2024] Efficient Streaming Language Models with Attention Sinks

Python 7,259 399 Updated Jul 11, 2024

The toolkit to test, validate, and evaluate your models and surface, curate, and prioritize the most valuable data for labeling.

Python 460 29 Updated May 23, 2025
Python 1,026 94 Updated Jan 4, 2024

Code for the AAAI 2024 Oral paper "OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models".

Python 72 8 Updated Mar 7, 2024

QLoRA: Efficient Finetuning of Quantized LLMs

Jupyter Notebook 10,986 877 Updated Jun 10, 2024