-
bytedance
- beijing-china
Stars
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
AI 时代的伯克希尔:基于 Claude Code / Codex 的价值投资研究框架。巴菲特·芒格·段永平·李录四大师方法论 + 多Agent并行研究。| AI-era Berkshire: a value investing research framework built for Claude Code / Codex. 4 masters' methodologies + multi…
AI Infra 全栈从0入门学习资料:https://caomaolufei.github.io/AIInfraGuide/
🤗 ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
This project aims to replicate mainstream open-source model architectures with limited computational resources, implementing mini models with 100-200M parameters.
把书、长视频、播客等高价值内容蒸馏成可执行的 Agent Skills
DuckDB is an analytical in-process SQL database management system
A composable and fully extensible C++ execution engine library for data management systems.
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
🎓从0开始训练一个大模型Minimind项目的超详细解析,包括但不限于用到的架构,算法,以及大模型面试经验
🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.
The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.
Implement a Pytorch-like DL library in C++ from scratch, step by step
hpc 教程,包含集合通信(mpi、nccl)、cuda 编程、向量化 SIMD、RDMA 通信等
Video+code lecture on building nanoGPT from scratch
LLM teach people something about LLM
Machine Learning Engineering Open Book
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
👀「大模型」2小时从0训练65M参数的视觉多模态VLM!Train a 65M-parameter VLM from scratch in just 2h!
🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
Paimon-cpp is a high-performance C++ implementation of Apache Paimon.
Sutskever 30 implementations inspired by https://papercode.vercel.app/ | For Agents, use https://github.com/pageman/Sutskever-Agent | Polyglot / Multi-Backed version at https://github.com/pageman/s…
An LLM training framework built from the ground up, featuring a custom BumbleBee architecture and end-to-end support for multiple open-source models across Pretraining → SFT → RLHF/DPO.
Algorithm powering the For You feed on X