Skip to content
View sh0416's full-sized avatar
🏃
🏃

Organizations

@PoApper @lmgsg-sh0416

Block or report sh0416

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Tools and prompt templates used to build and evaluate SWE-rebench-v2 tasks for the paper.

Python 78 6 Updated Mar 12, 2026

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,138 400 Updated Aug 6, 2026

A version of verl to support diverse tool use [TMLR 2026]

Python 1,029 88 Updated Jul 15, 2026

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,095 1,581 Updated Aug 9, 2026

Repository-level QA benchmark for software engineering LLMs

CSS 1 Updated May 15, 2026

Rethinking Code Editing for Efficient Software Engineering Agents

Python 12 1 Updated Jul 13, 2026
Python 57 8 Updated Apr 7, 2026

The agent that grows with you

Python 227,984 44,784 Updated Aug 10, 2026
Jupyter Notebook 1,438 211 Updated Dec 22, 2025

Agent framework and applications built upon Qwen>=3.0, featuring Function Calling, MCP, Code Interpreter, RAG, Chrome extension, etc.

Python 16,941 1,691 Updated Mar 4, 2026

Extracted system prompts from Anthropic - Claude Fable 5, Opus 5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-5.6-Sol, Codex. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Curso…

JavaScript 62,635 10,292 Updated Aug 7, 2026

👩‍⚖️ Agent-as-a-Judge: The Magic for Open-Endedness

HTML 806 106 Updated Mar 28, 2026

The first real AI developer

Python 33,710 3,479 Updated Jun 18, 2026

open source SWE-Atlas

Shell 66 5 Updated Jul 20, 2026

TDD-Bench-Verified is a new benchmark for generating test cases for test-driven development (TDD)

Python 34 6 Updated Jul 21, 2026

Exercism exercises in Python.

Python 2,472 1,517 Updated Aug 6, 2026

aider is AI pair programming in your terminal

Python 48,084 4,831 Updated May 22, 2026

Agentless🐱: an agentless approach to automatically solve software development problems

Python 2,093 237 Updated Dec 22, 2024

Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.

Python 16,773 1,237 Updated Mar 24, 2026

A conda-forge distribution.

Shell 10,076 524 Updated Aug 4, 2026

Supercharge Your LLM Application Evaluations 🚀

Python 15,238 1,611 Updated Feb 24, 2026

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

JavaScript 99,452 5,471 Updated Aug 7, 2026
Python 498 72 Updated Aug 4, 2026

PyTorch implementation of soft actor critic

Python 943 190 Updated Jul 17, 2025

An alignment auditing agent capable of quickly exploring alignment hypothesis

Python 1,283 211 Updated Jul 22, 2026

Realistic examples of building evals and optimizing agents with Harbor

Python 165 12 Updated Apr 23, 2026

Code for Paper: Training Software Engineering Agents and Verifiers with SWE-Gym [ICML 2025]

Jupyter Notebook 720 44 Updated Jul 29, 2025
Next