Lists (1)
Sort Name ascending (A-Z)
Stars
AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents
An Illusion of Progress? Assessing the Current State of Web Agents
[NeurIPS'23 Spotlight] "Mind2Web: Towards a Generalist Agent for the Web" -- the first LLM-based web agent and benchmark for generalist web agents
A QEC Evaluator built on Stim. Automated DEM construction.
Official Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
leeyeel / claude-code-sourcemap
Forked from Onewon/claude-codeclaude-code full original source code from source maps
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
An official implementation of DanceGRPO: Unleashing GRPO on Visual Generation
Enjoy the magic of Diffusion models!
DreamGen: Nvidia GEAR Lab's initiative to solve the robotics data problem using world models
🏘️ Scaling Embodied AI by Procedurally Generating Interactive 3D Houses
Official PyTorch implementation for ICML 2025 paper: UP-VLA.
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
openvla / openvla
Forked from TRI-ML/prismatic-vlmsOpenVLA: An open-source vision-language-action model for robotic manipulation.
robomimic: A Modular Framework for Robot Learning from Demonstration
RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicitation and Distillation Algorithms, and explore the Skill & V…
ALFRED - A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
slime is an LLM post-training framework for RL Scaling.
A mini-framework for running AI2-Thor with Docker.
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation (ICLR 2026)