Stars
Best practices for training DeepSeek, Mixtral, Qwen and other MoE models using Megatron Core.
The official repo of Pai-Megatron-Patch for LLM & VLM large scale training developed by Alibaba Cloud.
Awesome-LLM: a curated list of Large Language Model
An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related websites
A collection of various deep learning architectures, models, and tips