Highlights
- Pro
Lists (12)
Sort Name ascending (A-Z)
Stars
The official repo for "Vidi: Large Multimodal Models for Video Understanding and Editing"
OfficeCLI is the first and best Office suite purpose-built for AI agents to read, edit, and automate Word, Excel, and PowerPoint files. Free, open-source, single binary, no Office installation requ…
Skills for Designers and Engineers.
A curated list of AI PPT, PowerPoint automation, PPTX editing, and slide workflow tools.
Benchmarking Agentic Procedural 3D Modeling Via Code
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?
HowardLi1984 / OpenAI4S
Forked from PKU-YuanGroup/OpenAI4SClaude Science Spectra Part
Infinite Worlds with Versatile Interactions
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
TradingAgents: Multi-Agents LLM Financial Trading Framework
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
A modern GUI client based on Tauri, designed to run in Windows, macOS and Linux for tailored proxy experience
Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees.…
[ACL 2024 Findings] "TempCompass: Do Video LLMs Really Understand Videos?", Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, Lu Hou
Local AI filmmaking studio — skills, canvas, timeline — driven from your coding agent.
The code for "InstructSAM: Segment Any Instance with Any Instructions"
AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and supp…
[CVPR 2026] VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking
Official Python client library for the OpenReview API
Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
👾 E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding (NeurIPS 2024)
🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
A curated list of large VLM-based VLA models for robotic manipulation.
Cambrian-S: Towards Spatial Supersensing in Video
[NeurIPS 2025] Deep Memory Backtracking for Long Video Understanding