Starred repositories
AI agents running research on single-GPU nanochat training automatically
Static suckless single batch CUDA-only qwen3-0.6B mini inference engine
Examples for Recommenders - easy to train and deploy on accelerated infrastructure.
flash attention tutorial written in python, triton, cuda, cutlass
Efficient triton implementation of Native Sparse Attention.
Exploring Applications of GRPO
一个手把手教你从零开始编写GPT并训练大语言模型的教程
🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
Self-paced bootcamp on Generative AI. Tutorials on ML fundamentals, Ollama, LLMs, RAGs, LangChain, LangGraph, Fine-tuning, DSPy & AI Agents (CrewAI), (Using ChatGPT, gpt-oss, Claude, Qwen, Gemma, L…
This is a repository used by individuals to experiment and reproduce the pre-training process of LLM.
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), ga…
Machine Learning Journal for Intermediate to Advanced Topics.
test tensorrt c++/python api, c++/python plugins.
Writing an OS in 1,000 lines.
RAFT contains fundamental widely-used algorithms and primitives for machine learning and information retrieval. The algorithms are CUDA-accelerated and form building blocks for more easily writing …
[🔥updating ...] AI 自动量化交易机器人(完全本地部署) AI-powered Quantitative Investment Research Platform. 📃 online docs: https://ufund-me.github.io/Qbot ✨ :news: qbot-mini: https://github.com/Charmve/iQuant
搜索、推荐、广告、用增等工业界实践文章收集(来源:知乎、Datafuntalk、技术公众号)
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
how to optimize some algorithm in cuda.
deepspeedai / Megatron-DeepSpeed
Forked from NVIDIA/Megatron-LMOngoing research training transformer language models at scale, including: BERT & GPT-2
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
RPC framework based on C++ Workflow. Supports SRPC, Baidu bRPC, Tencent tRPC, thrift protocols.
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
Transformer Explained Visually: Learn How LLM Transformer Models Work with Interactive Visualization
Repository hosting code for "Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations" (https://arxiv.org/abs/2402.17152).
A high-throughput and memory-efficient inference and serving engine for LLMs