Skip to content
View nissymori's full-sized avatar

Block or report nissymori

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.

Python 4,019 294 Updated Aug 19, 2026

『因果AI ―コードファーストで学ぶ因果推論―』サポートページ

Jupyter Notebook 2 1 Updated Jul 23, 2026

Open-source Dreamer world-model implementation in JAX

Python 359 28 Updated Aug 5, 2026

CUDA Craftax-Classic: 7x faster RL training than JAX

Cuda 76 2 Updated Jul 25, 2026

Official implementation of DiscoGen, for "Procedural Generation of Algorithm Discovery Tasks in Machine Learning"

Python 50 10 Updated Aug 4, 2026

Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games

Python 4 Updated Jul 12, 2026

Open reimplementation of Sakana Fugu — the 'one model to command them all' LLM orchestrator. Read → run → train → serve.

Python 451 82 Updated Jun 22, 2026

[ICLR 2023] Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary Classification

Python 24 Updated Aug 18, 2026

CapCode: Detecting cheating in coding agents with capped, randomized tests

Python 5 Updated Jun 8, 2026

CapReward: Penalizing implausibly high pass rates to prevent reward hacking in coding RL

Python 5 Updated Jun 8, 2026

The Official JAX Code for "Retry Policy Gradients for Continuous Action Spaces"

Python 3 Updated Jun 5, 2026

Policy gradient reinforcement learning beyond the mean, e.g., Pass@k, Max@k, TopM@K, CVaR, VaR, Quantiles, Trimmed Means or other functionals of the reward distribution based on unbiased order-stat…

Python 13 1 Updated Jun 5, 2026

Proppo is a prototype Automatic Propagation software library, a generalization of Automatic Differentiation.

Python 9 Updated Nov 30, 2022

One unified CLI for headless coding agent execution 🤖

TypeScript 39 4 Updated Aug 14, 2026

[ICML 2026] CapBencher toolkit: Give your LLM benchmark a built-in alarm for leakage and gaming

Python 11 1 Updated May 29, 2026

[English/Japanese] A curated list of awesome online-prediction papers, libraries, and resources. Created and hosted by MIRU2025 Young Researchers Program group 5.

8 Updated Aug 19, 2025

[ICML2026] Official JAX code for Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

Python 16 Updated Jul 3, 2026

A Python tool that automatically cleans, completes, and standardizes BibTeX entries using LLMs and web search.

Python 187 7 Updated Jun 10, 2026

Transform arXiv papers into a single LaTeX source that can be used as a prompt for asking LLMs questions about the paper.

Python 166 10 Updated Jun 10, 2026

MCP server that uses arxiv-to-prompt to fetch and process arXiv LaTeX sources for precise interpretation of mathematical expressions in scientific papers.

Python 144 17 Updated Jul 30, 2026

Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞

Python 14,044 1,642 Updated Aug 19, 2026

AI agents running research on single-GPU nanochat training automatically

Python 94,164 13,321 Updated Mar 26, 2026

A Simple and Universal Swarm Intelligence Engine, Predicting Anything. 简洁通用的群体智能引擎,预测万物

Python 71,217 11,071 Updated Aug 17, 2026

Implementation for our paper "Gradient Regularization prevents Reward Hacking in RLHF and RLVR". Implemented TRL and for Huggingface Transformers

Python 12 Updated Feb 24, 2026

A fast and soft pattern search for trillion-scale corpora.

Python 239 11 Updated Feb 28, 2026
C++ 14 1 Updated Feb 18, 2026

https://mahjong.chingru.com - Japanese Mahjong Font (Riichi Mahjong) turns tile notation like 7m7m7m2p3p4p into inline-SVG mahjong hands, with an embeddable shields.io-style Image API.

TypeScript 31 2 Updated Jul 31, 2026

High-Performance Research Environment for Riichi Mahjong

Rust 74 19 Updated Aug 14, 2026
Next