Skip to content
View QiJune's full-sized avatar

Organizations

@PaddlePaddle

Block or report QiJune

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

From Automated Idea Factory to Realization

Shell 1,343 116 Updated Jul 18, 2026

FlashInfer: Kernel Library for LLM Serving

Python 6,049 1,209 Updated Jul 28, 2026

A throughput-oriented high-performance serving framework for LLMs

Jupyter Notebook 970 50 Updated Mar 29, 2026

Dynamic Memory Management for Serving LLMs without PagedAttention

C 506 42 Updated Jul 17, 2026

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Python 5,425 429 Updated Jul 26, 2026

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

Python 14,231 2,615 Updated Jul 28, 2026

Minimalist ML framework for Rust

Rust 20,762 1,677 Updated Jul 27, 2026

LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.

Python 4,197 346 Updated Jul 28, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 87,380 19,940 Updated Jul 28, 2026

Inference code for Llama models

Python 59,527 9,798 Updated Jan 26, 2025

SCQL (Secure Collaborative Query Language) is a system that allows multiple distrusting parties to run joint analysis without revealing their private data.

Go 182 72 Updated Mar 18, 2026

Running large language models on a single GPU for throughput-oriented scenarios.

Python 9,364 590 Updated Oct 28, 2024

Simple samples for TensorRT programming

Python 1,663 351 Updated Jul 21, 2026

Kernl lets you run PyTorch transformer models several times faster on GPU with a single line of code, and is designed to be easily hackable.

Jupyter Notebook 1,586 99 Updated Jan 28, 2026

🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.

Python 34,168 7,181 Updated Jul 28, 2026

Synthesizer for optimal collective communication algorithms

Python 126 29 Updated Apr 8, 2024

Repo for external large-scale work

Python 6,551 715 Updated Apr 27, 2024

Transformer related optimization, including BERT, GPT

C++ 6,448 936 Updated Mar 27, 2024
Python 2,976 340 Updated Jul 9, 2026

Development repository for the Triton language and compiler

MLIR 19,797 3,048 Updated Jul 28, 2026

Microsoft Collective Communication Library

C++ 394 35 Updated Sep 20, 2023

Large-scale model inference.

Python 629 84 Updated Sep 12, 2023

OneFlow is a deep learning framework designed to be user-friendly, scalable and efficient.

C++ 9,417 1,013 Updated Dec 4, 2025

A baseline repository of Auto-Parallelism in Training Neural Networks

Python 145 20 Updated Jun 25, 2022

XGo is a programming language that reads like plain English. But it's also incredibly powerful — it lets you leverage assets from C/C++, Go, Python, and JavaScript/TypeScript, creating a unified so…

Go 9,439 566 Updated Jul 23, 2026

Kubernetes-native Deep Learning Framework

Python 744 115 Updated Jan 26, 2024

Training and serving large-scale neural networks with auto parallelization.

Python 3,180 362 Updated Dec 9, 2023

Flexible and powerful tensor operations for readable and reliable code (for pytorch, jax, TF and others)

Python 9,561 400 Updated Jul 5, 2026
Next