Skip to content
View IsThatYou's full-sized avatar

Highlights

  • Pro

Block or report IsThatYou

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Automated auditing pipeline for LLM and agent benchmarks — surfaces task ambiguity, environment conflicts, and evaluation bugs.

HTML 13 1 Updated May 27, 2026

C++-based high-performance parallel environment execution engine (vectorized env) for general RL environments.

C++ 1,497 140 Updated Jul 17, 2026
JavaScript 11 1 Updated May 1, 2026

Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepowe…

TeX 11,818 862 Updated Jun 16, 2026

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…

Python 14,877 1,313 Updated Aug 18, 2026

A holistic framework for advancing LLMs as data science agents

Python 59 10 Updated May 19, 2026

GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's TerminalBench leaderboard.

Python 406 26 Updated Aug 24, 2025

One local control plane for every AI agent: route across models, fuse new capabilities, orchestrate tools, and stay fully in control.

TypeScript 36,721 3,070 Updated Aug 18, 2026

Claude Code. Any Model. The most powerful AI coding agent now speaks every language.

TypeScript 966 131 Updated Aug 19, 2026

🔥 Comprehensive survey on Context Engineering: from prompt engineering to production-grade AI systems. hundreds of papers, frameworks, and implementation guides for LLMs and AI agents.

3,278 276 Updated May 28, 2026

Examples of my Claude Code infrastructure with skill auto-activation, hooks, and agents

TypeScript 10,004 1,229 Updated Jul 13, 2026
Python 1,306 135 Updated May 20, 2026

An interface library for RL post training with environments.

Python 2,509 427 Updated Aug 18, 2026

PyTorch-native post-training at scale

Python 702 103 Updated Aug 18, 2026

Multiturn Meta-Bandit LLM RL Training

Python 11 Updated Oct 14, 2025

Efficient Triton Kernels for LLM Training

Python 6,572 583 Updated Aug 18, 2026

Democratizing Reinforcement Learning for LLMs

Python 5,790 607 Updated Aug 19, 2026

slime is an LLM post-training framework for RL Scaling.

Python 8,126 1,159 Updated Aug 16, 2026

Awesome curated collection of images and prompts generated by gemini-2.5-flash-image (aka Nano Banana) state-of-the-art image generation and editing model. Explore AI generated visuals created with…

JavaScript 8,813 891 Updated Sep 8, 2025

A curated list of awesome Deep Reinforcement Learning resources.

899 87 Updated Jul 13, 2025

[NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards

Python 1,487 123 Updated Apr 17, 2026

OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.

C++ 5,414 1,166 Updated Aug 12, 2026

This repository contains a curated collection of 300+ case studies from over 80 companies, detailing practical applications and insights into machine learning (ML) system design. The contents are o…

10,945 1,679 Updated Aug 5, 2025

Archon provides a modular framework for combining different inference-time techniques and LMs with just a JSON config file.

Python 211 31 Updated Mar 7, 2025

Super-Efficient RLHF Training of LLMs with Parameter Reallocation

Python 336 22 Updated Apr 24, 2025

A framework for few-shot evaluation of language models.

Python 13,711 3,498 Updated Aug 14, 2026

[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward

Python 958 78 Updated Feb 16, 2025

Together Mixture-Of-Agents (MoA) – 65.1% on AlpacaEval with OSS models

Python 2,966 385 Updated Jan 7, 2025

A quick guide (especially) for trending instruction finetuning datasets

3,411 234 Updated Nov 28, 2023

[ACL 2024] Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications

Python 20 2 Updated Apr 9, 2026
Next