-
BeingBeyond
- Beijing, China
-
23:36
(UTC +08:00) - https://zhangwp.com
- @zhang_wanpeng
- in/zawnpn
Highlights
Stars
Open source evals for physical AI. Run any LLM/VLA on any arm/humanoid against any real/sim benchmark.
This repository provides the BeingBeyond D1 education-version SDK, example Python scripts, and basic guidance for environment setup and common first-time troubleshooting.
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models (ECCV 2026)
Lean 4 programming language and theorem prover
A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
An open-source, GPU-accelerated physics simulation engine built upon NVIDIA Warp, specifically targeting roboticists and simulation researchers.
Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
Conservative Offline Robot Policy Learning via Posterior-Transition Reweighting
Agent skill that removes signs of AI-generated writing from text
High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.
Programmatic access to Gemini Notebook - via command-line interface (CLI), Model Context Protocol (MCP) server, and AI agent skills.
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild (CVPR 2026)
🔥 LeetCode for PyTorch — practice implementing softmax, attention, GPT-2 and more from scratch with instant auto-grading. Jupyter-based, self-hosted or try online.
A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modular configuration across SFT, RLVR, and evaluation workflows.
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
[RSS 2026] Causal video-action world model for generalist robot control
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos (ICML 2026)
UniTacHand: Unified Spatio-Tactile Representation for Human-to-Dexterous-Hand Skill Transfer
Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos (CVPR 2026)
An implementation of a data collector app.
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models (ECCV 2026)
This repository provides the BeingBeyond D1 SDK, example Python scripts, and basic guidance for environment setup and common first-time troubleshooting.
This is the official repo for the paper "LongCat-Flash-Omni Technical Report"
OpenMMEgo: Enhancing Egocentric Understanding for LMMs with Open Weights and Data (NeurIPS 2025)