Stars
Code for URBER: Ultrafast Rule-Based Escape Routing Method for Large-Scale Sample Delivery Biochips
Phi-Bench (Φ-Bench): 85 open-source LLM-infrastructure engineering tasks for frontier LLMs & coding agents (KFC/LH/E2E). Self-contained public Dockerfile + offline scoring + reference solution per …
A high-efficiency Flash model for real-world agents.
A lightweight, AI-native training framework for large language models. Designed for fast iteration, reproducible experiments, and modular configuration across SFT, RLVR, and evaluation workflows.
💻 SETA: Scaling Environments for Terminal Agents - Environments
Fast, Sharp & Reliable Agentic Intelligence
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Step3-VL-10B: A compact yet frontier multimodal model achieving SOTA performance at the 10B scale, matching open-source models 10-20x its size.
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
STEP-GUI: The top GUI agent solution in the galaxy. Developed by the StepFun-GELab team and powered by StepFun’s cutting-edge research capabilities.
[ECCV 2026] DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
Dexbotic: Open-Source Vision-Language-Action Toolbox
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
🙌 OpenHands: AI-Driven Development
A benchmark for LLMs on complicated tasks in the terminal
The Open Cookbook for Top-Tier Code Large Language Model
StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
[🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s Multimodal Intelligence team.
[NeurIPS 2025] The official repository for our paper, "Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning".
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Large language models designed for formal theorem proving through tool-integrated reasoning.
A lightweight reinforcement learning framework that integrates seamlessly into your codebase, empowering developers to focus on algorithms with minimal intrusion.
DDN: A novel generative model with simple principles and unique properties. (ICLR 2025)