Stars
分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等
An LLM post-training framework with vLLM for RL Scaling
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
A simple calculation for LLM MFU.
Coding the transformer model architecture in one page for tutorial.
Langchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM, Qwen and…
NVIDIA Linux open GPU kernel module source
Swift Core ML 3 implementations of GPT-2, DistilGPT-2, BERT, and DistilBERT for Question answering. Other Transformers coming soon!
Samples for CUDA Developers which demonstrates features in CUDA Toolkit
Making large AI models cheaper, faster and more accessible
Shows how to check if a GPU is an Enterprise/Quadro GPU using NVML.
Learning Vim and Vimscript doesn't have to be hard. This is the guide that you're looking for 📖
🏆 A ranked list of awesome machine learning Python libraries. Updated weekly.
955 不加班的公司名单 - 工作 955,work–life balance (工作与生活的平衡)
A C++ implementation of the scalar-valued autograd engine micrograd
You like pytorch? You like micrograd? You love tinygrad! ❤️
A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like API
It is implementation of Research paper "DEEP GRADIENT COMPRESSION: REDUCING THE COMMUNICATION BANDWIDTH FOR DISTRIBUTED TRAINING". Deep gradient compression is a technique by which the gradients ar…
Open-source implementation of Google Vizier for hyper parameters tuning
text detection mainly based on ctpn model in tensorflow, id card detect, connectionist text proposal network