-
University of California, San Diego
- La Jolla, California
- zilongwang.me
Stars
Unofficial source-oriented reconstruction and extension of Grok Bot 0.18.0 for macOS
AgentENV (AENV) is a distributed platform for running agent environments at scale.
Open-source AI penetration testing tool to find and fix your app’s vulnerabilities.
SecCodeBench is a benchmark suite focusing on evaluating the security of code generated by large language models (LLMs).
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
The AI that really does things. Any OS. Any Platform. The lobster way. 🦞
An agent framework for building and evaluating general digital agents.
A lightweight, powerful framework for multi-agent workflows
PatchEval: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
All-in-One Sandbox for AI Agents that combines Browser, Shell, File, MCP and VSCode Server in a single Docker container.
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Official Implementation of "Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought"
Post-training with Tinker
Official implementation of "Flow Based Policy for Online Reinforcement Learning"
SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can also be employed for offensive cybersecurity or competitive coding challenges. [NeurIPS 2024]
[NeurIPS 2025 D&B Spotlight] Scaling Data for SWE-agents
Official implementation of paper "Learning to Optimize Multi-objective Alignment Through Dynamic Reward Weighting"
OSS-Fuzz - continuous fuzzing for open source software.
Sandboxed code execution for AI agents, locally or on the cloud. Massively parallel, easy to extend. Powering SWE-agent and more.
SkyRL: A Modular Full-stack RL Library for LLMs
[ICLR 2026] End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
"DeepCode: Open Agentic Coding (Agent Harness & Loop Engineering & Multi-Agent Orchestration)"
The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.
Implementation for FP8/INT8 Rollout for RL training without performence drop.