Stars
A curated list of practical Codex skills for automating workflows across the Codex CLI and API.
[ACL 2025] Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
This repository is established to store personal notes and annotated papers during daily research.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
🍁Vim cheat sheet with everything you want to know.
Aims to implement dual-port and multi-qp solutions in deepEP ibrc transport
Awesome list for LLM quantization
NVIDIA's launch, startup, and logging scripts used by our MLPerf Training and HPC submissions
Efficient Mixture of Experts for LLM Paper List
Language model alignment-focused deep learning curriculum
Tutel MoE: Optimized Mixture-of-Experts Library, Support GptOss/DeepSeek/Kimi-K2/Qwen3 using FP8/NVFP4/MXFP4
Official inference framework for 1-bit LLMs
主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题
Efficient Triton Kernels for LLM Training
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translatio…
基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型
[DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror
This repository contains the experimental PyTorch native float8 training UX
A high-throughput and memory-efficient inference and serving engine for LLMs
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance…
A curated collection of papers, tutorials, videos, and other valuable resources related to Mamba.
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
[TMLR 2024] Efficient Large Language Models: A Survey
Talk to any LLM with hands-free voice interaction, voice interruption, and Live2D taking face running locally across platforms
You like pytorch? You like micrograd? You love tinygrad! ❤️