-
Tsinghua University
- https://scholar.google.com/citations?hl=zh-CN&user=kMui170AAAAJ
Stars
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
FlashKDA: high-performance Kimi Delta Attention kernels
AgentENV (AENV) is a distributed platform for running agent environments at scale.
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
A generalist video MLLM built for fine-grained motion, long-form reasoning, temporal grounding, and online proactive response.
[ECCV2026] ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
Measuring frontier coding agents on original, long-horizon engineering tasks
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence
From Vision-Language-Action Models to a Real-World Robot Learning Stack
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
Kimi Code CLI — The Starting Point for Next-Gen Agents
A collection of skills for AI financial analysis.
My learning notes for ML SYS.
[ECCV 2026] Official code of GEM: Generative Supervision Helps Embodied Intelligence
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
Skill package for ML/CV/NLP paper writing, curated and adapted from Prof. Peng Sida's open notes for Codex, Claude Code, and Gemini.
Can Language Models Rebuild Programs From Scratch?
Beyond SFT-to-RL: Pre-alignment via Black-BoxOn-Policy Distillation for Multimodal RL
A benchmark for evaluating LLMs on Chinese traditional fortune telling — Bazi (八字) and Ziwei Doushu (紫微斗数).
📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
Extracted system prompts from Anthropic - Claude Fable 5, Opus 5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-5.6-Sol, Codex. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Curso…
SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
Reference code for the Meta-Harness paper.
Terrarium: Multi-turn data engine for evaluating and optimizing LLM agents in living environments.
🦞 ClawMark: A Living-World Benchmark for Multi-Day, Multimodal Coworker Agents
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
The agent that grows with you