Skip to content
View kaiyux's full-sized avatar
🎯
Focusing
🎯
Focusing
  • Beijing, China
  • 13:12 (UTC +08:00)

Organizations

@NVIDIA

Block or report kaiyux

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

The best-benchmarked open-source AI memory system. And it's free.

Python 57,775 7,440 Updated Jul 26, 2026

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

JavaScript 233,728 35,631 Updated Jul 27, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 384,254 80,730 Updated Jul 27, 2026

MLX: An array framework for Apple silicon

C++ 27,721 2,057 Updated Jul 25, 2026

Reference implementations of MLPerf® inference benchmarks

Python 1,604 641 Updated Jul 24, 2026

The Triton TensorRT-LLM Backend

939 141 Updated Jul 22, 2026

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

Python 14,223 2,609 Updated Jul 27, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 87,262 19,903 Updated Jul 27, 2026

天涯 kkndme 神贴聊房价

19,420 3,852 Updated Jun 4, 2026

Tensor library for machine learning

C++ 15,062 1,742 Updated Jul 17, 2026

AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.

Python 185,705 46,068 Updated Jul 26, 2026

Code for loralib, an implementation of "LoRA: Low-Rank Adaptation of Large Language Models"

Python 13,689 918 Updated Dec 17, 2024

Stable Diffusion with Core ML on Apple Silicon

Python 17,949 1,071 Updated Jul 3, 2025

A Python framework for GPU-accelerated simulation, robotics, and machine learning.

Python 6,905 571 Updated Jul 27, 2026

PyTriton is a Flask/FastAPI-like interface that simplifies Triton's deployment in Python environments.

Python 846 61 Updated Aug 13, 2025

Open source cross-platform compiler for compute-intensive loops used in AI algorithms, from Microsoft Research

C++ 115 21 Updated Oct 10, 2023

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

Python 19,022 3,034 Updated Apr 14, 2026

LLM inference in C/C++

C++ 121,703 21,027 Updated Jul 27, 2026

Inference code for Llama models

Python 59,527 9,798 Updated Jan 26, 2025

Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.

Python 31,251 3,767 Updated Jul 26, 2026

Running large language models on a single GPU for throughput-oriented scenarios.

Python 9,364 590 Updated Oct 28, 2024

Container plugin for Slurm Workload Manager

C 454 46 Updated May 12, 2026

Pipeline Parallelism for PyTorch

Python 786 88 Updated Aug 21, 2024

Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture. Basically ChatGPT but with PaLM

Python 7,867 674 Updated May 29, 2026

A latent text-to-image diffusion model

Jupyter Notebook 73,226 10,581 Updated Jun 18, 2024

Examples demonstrating available options to program multiple GPUs in a single node or a cluster

Cuda 909 153 Updated Sep 26, 2025

AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (NVIDIA GPU) and MatrixCore (AMD GPU) inference.

Python 4,727 390 Updated Jul 14, 2026

Common source, scripts and utilities for creating Triton backends.

C++ 377 112 Updated Jul 22, 2026
Next