Stars
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
Offline optimization of your disaggregated Dynamo graph
High-Performance Sparse Linear Algebra on HBM-Equipped FPGAs Using HLS
Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling
Ongoing research training transformer models at scale
A Datacenter Scale Distributed Inference Serving Framework
”数学不难“ 之 《线性代数不难》上下册,66话题完册;欢迎批评指正
inference cookbook / inference 框架原理解析
Skills for Real Engineers. Straight from my .agents directory.
Create beautiful slides on the web using a coding agent's frontend skills
EOS is a dual-core operating system designed specifically for embodied intelligence, suitable for robots, drones, satellites or other scenarios requiring real-time and general capabilities.
Cross-platform, C implementation of the IETF QUIC protocol, exposed to C, C++, C# and Rust.
Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the effective training time by minimizing the downtime due to fa…
NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated computing environments
Venus Collective Communication Library, supported by SII and Infrawaves.
AIInfra(AI 基础设施)指AI系统从底层芯片等硬件,到上层软件栈支持AI大模型训练和推理。
Community maintained hardware plugin for vLLM on MetaX GPU
A Cloud Native Batch System (Project under CNCF)