Skip to content
View pzhao1799's full-sized avatar

Block or report pzhao1799

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.

Go 7,320 1,046 Updated Aug 14, 2026

Slurm on Kubernetes - a cluster management tool

Go 11 4 Updated Jan 15, 2026

cuTile is a programming model for writing parallel kernels for NVIDIA GPUs

Python 2,126 143 Updated Aug 14, 2026

A framework for efficient model inference with omni-modality models

Python 6,115 1,468 Updated Aug 14, 2026

Gateway API Inference Extension

Go 740 307 Updated Aug 12, 2026

Achieve state of the art inference performance with modern accelerators on Kubernetes

Shell 4,028 679 Updated Aug 14, 2026

A Datacenter Scale Distributed Inference Serving Framework

Rust 7,767 1,441 Updated Aug 14, 2026

An AI Hedge Fund Team

Python 62,850 11,064 Updated Aug 7, 2026

An open-source, privacy-first, self-hosted knowledge workspace where humans and AI agents work together 开源、隐私优先、自托管的知识工作空间,让人与智能体在此协作

TypeScript 45,800 2,950 Updated Aug 14, 2026

Cost-efficient and pluggable Infrastructure components for GenAI inference

Go 5,009 646 Updated Aug 14, 2026

Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation

8,044 294 Updated May 15, 2025

AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.

Go 1,242 132 Updated Jul 31, 2026

[WIP] Resources for AI engineers. Also contains supporting materials for the book AI Engineering (Chip Huyen, 2025)

Jupyter Notebook 17,026 2,484 Updated Jul 3, 2026

Optimized Agentic and LLM Bulk Processing Over Your Data

Python 1,659 150 Updated Jul 3, 2026

A heap memory profiler for Linux

C++ 4,148 241 Updated Aug 12, 2026

High performance self-hosted photo and video management solution.

TypeScript 110,530 6,529 Updated Aug 14, 2026

A playbook for effectively prompting post-trained LLMs

901 38 Updated Jan 21, 2025

Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.

Python 71,477 6,447 Updated Aug 14, 2026

Fast Open-Source Search & Clustering engine × for Vectors & Arbitrary Objects × in C++, C, Python, JavaScript, Rust, Java, Objective-C, Swift, C#, GoLang, and Wolfram 🔍

C++ 4,264 340 Updated Jul 10, 2026

AI Observability & Evaluation

Python 11,055 1,057 Updated Aug 14, 2026

Adding guardrails to large language models.

Python 7,285 672 Updated Aug 14, 2026

DSPy: The framework for programming—not prompting—language models

Python 37,187 3,215 Updated Aug 14, 2026

An enterprise-grade AI retriever designed to streamline AI integration into your applications, ensuring cutting-edge accuracy.

Python 295 41 Updated Jun 26, 2025

LLM inference in C/C++

C++ 123,930 21,711 Updated Aug 14, 2026

Tensor library for machine learning

C++ 15,168 1,776 Updated Aug 14, 2026

[ACL'25] Official Code for LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs

Python 317 29 Updated Jul 13, 2025

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 31,816 7,888 Updated Aug 14, 2026

A simple RPC framework with protobuf service definitions

Go 7,527 325 Updated Aug 5, 2024

Open source, local, and self-hosted highly optimized language inference server supporting ASR/STT, TTS, and LLM across WebRTC, REST, and WS

Python 512 61 Updated Feb 12, 2026
Next