A clean implementation based on AlphaZero for any game in any framework + tutorial + Othello/Gobang/TicTacToe/Connect4 and more
-
Updated
Jan 1, 2025 - Jupyter Notebook
A clean implementation based on AlphaZero for any game in any framework + tutorial + Othello/Gobang/TicTacToe/Connect4 and more
OpenDILab Decision AI Engine. The Most Comprehensive Reinforcement Learning Framework B.P.
[NeurIPS 2023 Spotlight] LightZero: A Unified Benchmark for Monte Carlo Tree Search in General Sequential Decision Scenarios (awesome MCTS)
An artificial intelligence platform for the StarCraft II with large-scale distributed training and grand-master agents.
The official implementation of Self-Play Fine-Tuning (SPIN)
The official implementation of Self-Play Preference Optimization (SPPO)
A Massively Parallel Large Scale Self-Play Framework
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
Train a neural network to PvP in Old School RuneScape using reinforcement learning.
🏆 Threes! AI: deck-aware expectimax, N-tuple TD learning, and AlphaZero
A custom MARL (multi-agent reinforcement learning) environment where multiple agents trade against one another (self-play) in a zero-sum continuous double auction. Ray [RLlib] is used for training.
Search Self-Play: Pushing the Frontier of Agent Capability without Supervision
AlphaZero implementation for Othello, Connect-Four and Tic-Tac-Toe based on "Mastering the game of Go without human knowledge" and "Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm" by DeepMind.
A very fast implementation of AlphaZero, applied to games like Splendor, Santorini, The Little Prince, … Browser version available
Offline Clash Royale AI research: native simulation, imitation learning, PPO, and interactive evaluation. 离线对局模拟、模仿学习、强化学习与交互评测。
Backgammon OpenAI Gym
The exact codes used by the team "liveinparis" at the kaggle football competition ranked 6th/1141
[ICLR'26] MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
TD-Gammon implementation
CFR 德州扑克 AI 源码,面向 HUNL 策略训练、反事实价值、自我博弈与评估研究。C++/Python poker AI research system.
To associate your repository with the self-play topic, visit your repo's landing page and select "manage topics."