Skip to content
View lambda7xx's full-sized avatar
  • Shanghai Jiao Tong University
  • Shanghai

Block or report lambda7xx

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

DeepSeek-native AI coding agent for your terminal. Engineered around prefix-cache stability — leave it running.

Go 32,916 2,122 Updated Aug 8, 2026

Mixture-of-experts (MoE) training megakernel for NVL72s

Python 471 46 Updated Aug 5, 2026
C++ 6 1 Updated Jul 24, 2026

An efficient service for transparent GPU multiplexing with VRAM oversubscription

Rust 11 5 Updated Jun 28, 2026

C++ IPC Library

C++ 3 Updated May 11, 2026
Python 2 Updated Mar 6, 2026

Embodied AI Operating System (EAIOS)

Rust 313 45 Updated Aug 7, 2026
88 Updated Jul 17, 2026

Flash Vision-Language-Action Inference for Autonomous Driving

Python 59 7 Updated Jul 6, 2026

分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等

Jupyter Notebook 3,398 324 Updated Aug 7, 2026

Machine Learning Engineering Open Book

Python 18,530 1,192 Updated Aug 8, 2026

北京大学未名超算队与北京大学学生 Linux 俱乐部合办的暑期 AI Infra 系列活动仓库

Cuda 128 91 Updated Jul 30, 2026

Open Frontier Intelligence

8,190 622 Updated Aug 6, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 1,045 111 Updated Aug 7, 2026

Holistic evaluation of AI serving systems

Python 7 6 Updated Aug 3, 2026

A distributed framework for LLM agents

Python 607 26 Updated Aug 6, 2026

Production-ready MoE load balancing via real-time expert replication

Cuda 223 9 Updated Jul 17, 2026

A proactive on-device intelligent operating system for HarmonyOS, featuring proactive perception, autonomous decision-making, lightweight user modeling, and local LLM deployment across smartphones …

Python 7 Updated Jul 17, 2026

An open-source toolkit for BigMac-style pipeline-parallel training of multimodal large language models.

Python 32 4 Updated Jul 21, 2026

High-performance single-GPU inference for selected model checkpoints and GPUs.

C++ 276 39 Updated Jul 31, 2026

PhyAI is a high-performance framework for running Physical AI models (VLA, WAM, and beyond), supporting both cloud-based serving and on-device deployment.

Python 75 19 Updated Aug 2, 2026

Inspect: A framework for large language model evaluations

Python 2,497 634 Updated Aug 8, 2026

Implementation of plug in and play Attention from "LongNet: Scaling Transformers to 1,000,000,000 Tokens"

Python 726 60 Updated Jan 7, 2024
Python 6 1 Updated Nov 16, 2025

Fast and memory-efficient classical machine learning operators

Python 555 45 Updated Aug 4, 2026

High-Throughput Batch Inference

C++ 13 1 Updated Aug 6, 2026

TurboServe: Serving Streaming Video Generation Efficiently and Economically

Python 40 3 Updated Jul 12, 2026
Python 890 84 Updated Aug 6, 2026
Next