-
Tencent
- Shenzhen, China
Stars
[Official Repo] JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
Code for MIRA: Multiplayer Interactive World Models with Representation Autoencoders
OpenRW "Open ReWrite" is an un-official open source recreation of the classic Grand Theft Auto III game executable
A curated list of streaming agents, covering AI systems (LLM/VLM/VLA/video gen.) that are time-sensitive, real-time responsive, proactive, and capable of processing unbounded streaming inputs.
🎮 A curated list of awesome game datasets, and tools to artificial intelligence in games
Official code for "Stateful Visual Encoders for Vision-Language Models"
[ICLR 2025] LAPA: Latent Action Pretraining from Videos
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
Official code base for LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
[ICLR 2026] LLM/VLM gaming agents and model evaluation through games.
A Gym-like environment for Reinforcement Learning in Rocket League
Agentic Hours-Long Video Editing via Music Synchronization
Benchmark environment for evaluating vision-language models (VLMs) on popular video games!
Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with C…
"CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
A generative game world for your OpenClaw: turn a single sentence into an endless adventure with dynamic narratives, rendered scenes, and RPG mechanics. Play manually, or let guest agents take over.
[CVPR 2026] Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
Official Repo for paper: Scaling Behavior Cloning Improves Causal Reasoning: An Open Model for Real-Time Video Game Playing
Algorithm powering the For You feed on X
openvla / openvla
Forked from TRI-ML/prismatic-vlmsOpenVLA: An open-source vision-language-action model for robotic manipulation.
[CVPR 2026] SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time
Googles NotebookLM but local
[CVPR 2026] PersonaLive! : Expressive Portrait Image Animation for Live Streaming
Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries