Lists (1)
Sort Name ascending (A-Z)
Starred repositories
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
MultiModal Audio Generation in Raw Waveform Space.
VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
PyTorch implementation of JiT https://arxiv.org/abs/2511.13720
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Give your agents the power of the Hugging Face ecosystem
Fast Multimodal Semantic Deduplication & Filtering
AI+金融(量化):1.多因子股票量化框架开源教程 2.学界和业界的经典资料收录 3.AI + 金融的相关工作,包括LLM, Agent, benchmark(evaluation), etc.
A unified architecture deep learning framework designed specifically for ultra-large-scale sparse models.
Algorithm powering the For You feed on X
An Open Foundation Model and Benchmark to Accelerate Generative Recommendation
🔥 OneThinker: All-in-one Reasoning Model for Image and Video [CVPR 2026]
[ECCV 2026] Glance: Accelerating Diffusion Models with 1 Sample
A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing
alpha101, alpha191, alphalens, backtrader, 量化研究
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
What Is a Good Caption? A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness