-
Researcher at Tencent AI Lab
- China
Lists (1)
Sort Name ascending (A-Z)
Stars
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
Official JAX implementation of End-to-End Test-Time Training for Long Context
A Foundation Model for Generalist Gaming Agents
One portable memory layer for every AI agent: local-first, Markdown-native, user-owned, and self-evolving across apps, tools, and workflows.
Unofficial implementation of Titans, SOTA memory for transformers, in Pytorch
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
Codebase of 'From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model'
Marco Search Agent for Realistic and Challenging Agentic Search
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
Mamba-Chat: A chat LLM based on the state-space model architecture 🐍
A highly capable 2.4B lightweight LLM using only 1T pre-training data with all details.
RWKV-LM-V7(https://github.com/BlinkDL/RWKV-LM) Under Lightning Framework
Unofficial PyTorch implementation of Attention Free Transformer (AFT) layers by Apple Inc.
Structured state space sequence models
Official PyTorch implementation of One-Minute Video Generation with Test-Time Training
This repo contains the source code for RULER: What’s the Real Context Size of Your Long-Context Language Models?
Doing simple retrieval from LLM models at various context lengths to measure accuracy
ChatRWKV is like ChatGPT but powered by RWKV (100% RNN) language model, and open source.