Stars
An open ecosystem of parametric human models and perception stacks, starting with GNM Head.
This is the official repository for the paper: A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models (ICLR 2026).
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
PaperBanana: Automating Academic Illustration For AI Scientists
Official Code Repository for the ICLR 2026 Paper "Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding" & ICML 2026 Paper "Omnisapiens: A foundation model for …
[CVPR 2026] The released code of PerformRecast paper
[CVPR 2023] Code for "Learning Emotion Representations from Verbal and Nonverbal Communication"
A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like API
ARTalk generates realistic 3D head motions (lip sync, blinking, expressions, head poses) from audio in ⚡ real-time ⚡.
Analysis of EEG Signals and Facial Expressions for Continuous Emotion Detection code
Talk to any LLM with hands-free voice interaction, voice interruption, and Live2D taking face running locally across platforms
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Wan: Open and Advanced Large-Scale Video Generative Models
A complete head tracking pipeline from videos to NeRF/3DGS-ready datasets.
Accurate 3D Face Reconstruction with Weakly-Supervised Learning: From Single Image to Image Set (CVPRW 2019)
Summary of publicly available ressources such as code, datasets, and scientific papers for the FLAME 3D head model
[Official Code] Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction
Official implementation of AsymFlow, pi-Flow, GMFlow
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
ICT's Vision and Graphics Lab's morphable face model and toolkit
Efficient face emotion recognition in photos and videos
Retinaface get 80.99% in widerface hard val using mobilenet0.25.
[AAAI'21] Robust Lightweight Facial Expression Recognition Network with Label Distribution Training
[AAAI 2025] VQTalker: Towards Multilingual Talking Avatars through Facial Motion Tokenization
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
An open-source AI agent that brings the power of Gemini directly into your terminal.
The official implementation of 3DDFA_V3 in CVPR2024 (Highlight).