-
Shanghai Jiao Tong University
- Shanghai, China
Highlights
- Pro
Stars
Open-source framework for the research and development of foundation models.
Train transformer language models with reinforcement learning.
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
Minimalistic 4D-parallelism distributed training framework for education purpose
minted is a LaTeX package that provides syntax highlighting using the Pygments library. Highlighted source code can be customized using fancyvrb.
My learning notes for ML SYS.
NVSHMEM‑Tutorial: Build a DeepEP‑like GPU Buffer
A high-throughput and memory-efficient inference and serving engine for LLMs
Distributed Compiler and Optimized Parallel Kernels
Training materials associated with NVIDIA's CUDA Training Series (www.olcf.ornl.gov/cuda-training-series/)
DeepEP: an efficient expert-parallel communication library
Optimized primitives for collective multi-GPU communication
FlashInfer: Kernel Library for LLM Serving
The world’s fastest framework for building websites.
A fast communication-overlapping library for tensor/expert parallelism on GPUs.
2021年最新整理, C++ 学习资料,含C++ 11 / 14 / 17 / 20 / 23 新特性、入门教程、推荐书籍、优质文章、学习笔记、教学视频等
High performance Transformer implementation in C++.
A framework for writing fast and performant SQLite extensions in Rust
SJTU Canvas Helper——帮助您更快速便捷地使用上海交通大学课程平台。
《Hello 算法》:动画图解、一键运行的数据结构与算法教程。支持简中、繁中、English、日本語,提供 Python, Java, C++, C, C#, JS, Go, Swift, Rust, Ruby, Kotlin, TS, Dart 等代码实现
A library for building fast, reliable and evolvable network services.
Implementation of 💍 Ring Attention, from Liu et al. at Berkeley AI, in Pytorch
LlamaIndex is the document processing platform for AI