- Beijing, China
- https://aberhu.github.io/
Stars
[SIGGRAPH‘2026] PEAR :Pixel-aligned Expressive humAn mesh Recovery
Academic Research Skills for Claude Code: research → write → review → revise → finalize
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Models
from vibe coding to agentic engineering - practice makes claude perfect
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
同事.skill、老板.skill、前任.skill、自己.skill、永生.skill、女娲.skill……
Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
Evaluate and improve models and agents using environments
Scalable toolkit for efficient model reinforcement
VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
LEAKED SYSTEM PROMPTS FOR CHATGPT, CLAUDE, GEMINI, GROK, PERPLEXITY, CURSOR, LOVABLE, REPLIT, AND MORE! - AI SYSTEMS TRANSPARENCY FOR ALL! 👐
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
"AI-Trader: 100% Fully-Automated Agent-Native Trading"
Mobile-Agent: The Powerful GUI Agent Family
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
A comprehensive list of papers for the definition of World Models and using World Models for General Video Generation, Embodied AI, and Autonomous Driving, including papers, codes, and related webs…
The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.
Fully Open Framework for Democratized Multimodal Training
[CVPR 2026] SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
A Survey of Reinforcement Learning for Large Reasoning Models
Unlimited-length talking video generation that supports image-to-video and video-to-video generation
🦛 CHONK docs with Chonkie ✨ — The lightweight ingestion library for fast, efficient and robust RAG pipelines
Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
E2M converts various file types (doc, docx, epub, html, htm, url, pdf, ppt, pptx, mp3, m4a) into Markdown. It’s easy to install, with dedicated parsers and converters, supporting custom configs. E2…
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]
Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’