-
The University of Tokyo
- Tokyo, Japan
-
21:36
(UTC -12:00) - https://nissymori.github.io/
- @nissymori1
Stars
Open-source Dreamer world-model implementation in JAX
CUDA Craftax-Classic: 7x faster RL training than JAX
Official implementation of DiscoGen, for "Procedural Generation of Algorithm Discovery Tasks in Machine Learning"
Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games
Open reimplementation of Sakana Fugu — the 'one model to command them all' LLM orchestrator. Read → run → train → serve.
[ICLR 2023] Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary Classification
CapCode: Detecting cheating in coding agents with capped, randomized tests
CapReward: Penalizing implausibly high pass rates to prevent reward hacking in coding RL
The Official JAX Code for "Retry Policy Gradients for Continuous Action Spaces"
Policy gradient reinforcement learning beyond the mean, e.g., Pass@k, Max@k, TopM@K, CVaR, VaR, Quantiles, Trimmed Means or other functionals of the reward distribution based on unbiased order-stat…
Proppo is a prototype Automatic Propagation software library, a generalization of Automatic Differentiation.
One unified CLI for headless coding agent execution 🤖
[ICML 2026] CapBencher toolkit: Give your LLM benchmark a built-in alarm for leakage and gaming
[English/Japanese] A curated list of awesome online-prediction papers, libraries, and resources. Created and hosted by MIRU2025 Young Researchers Program group 5.
[ICML2026] Official JAX code for Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
A Python tool that automatically cleans, completes, and standardizes BibTeX entries using LLMs and web search.
Transform arXiv papers into a single LaTeX source that can be used as a prompt for asking LLMs questions about the paper.
MCP server that uses arxiv-to-prompt to fetch and process arXiv LaTeX sources for precise interpretation of mathematical expressions in scientific papers.
Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
AI agents running research on single-GPU nanochat training automatically
A Simple and Universal Swarm Intelligence Engine, Predicting Anything. 简洁通用的群体智能引擎,预测万物
Implementation for our paper "Gradient Regularization prevents Reward Hacking in RLHF and RLVR". Implemented TRL and for Huggingface Transformers
A fast and soft pattern search for trillion-scale corpora.
https://mahjong.chingru.com - Japanese Mahjong Font (Riichi Mahjong) turns tile notation like 7m7m7m2p3p4p into inline-SVG mahjong hands, with an embeddable shields.io-style Image API.
High-Performance Research Environment for Riichi Mahjong
Open Bandit Pipeline: a python library for bandit algorithms and off-policy evaluation