Skip to content
View RS2002's full-sized avatar
🎸
🎸

Block or report RS2002

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Skill package for ML/CV/NLP paper writing, curated and adapted from Prof. Peng Sida's open notes for Codex, Claude Code, and Gemini.

5,927 290 Updated Jun 23, 2026

Code for tasks on Cainiao-LaDe (Last-mile Delivery dataset).

Python 107 31 Updated Dec 23, 2025
Python 12 4 Updated Apr 10, 2025
Python 11 1 Updated Apr 10, 2024
Python 3 Updated Feb 22, 2026

Implementation for “Hierarchical Optimization via LLM-Guided Objective Evolution for Mobility-on-Demand Systems.” (NeurIPS 2025)

Python 4 2 Updated Oct 24, 2025

xingtian is a componentized library for the development and verification of reinforcement learning algorithms

Python 318 89 Updated Sep 12, 2023

A next-generation LLM4AD platform focused on intuitive UI interactions and seamless collaboration with AI agents, making automated algorithm design more accessible and easier to use

Python 380 26 Updated Aug 10, 2026
Python 22 7 Updated Jun 23, 2026

Official Implementation of wd1

Python 32 1 Updated Sep 25, 2025

Code for paper "SPG Sandwiched Policy Gradient for Masked Diffusion Language Models"

Python 63 7 Updated Oct 29, 2025

Mean Field Multi-Agent Reinforcement Learning

Python 422 105 Updated Mar 11, 2020

Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"

Python 1,070 140 Updated May 30, 2026

To make music production easier, we introduce Amadeus , a novel MIDI generation framework. While significantly improving generation quality, we have achieved a speedup of at least 4x compared to pu…

Python 17 Updated Aug 29, 2025

Official implementation of "Diffusion Language Models Know the Answer Before Decoding"

Jupyter Notebook 60 1 Updated Apr 28, 2026

A PyTorch library for all things Reinforcement Learning (RL) for Combinatorial Optimization (CO)

Python 897 152 Updated May 12, 2026

[EMNLP 2024 (main)] Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters

Python 14 2 Updated Nov 5, 2024

BertViz: Visualize Attention in Transformer Models

Python 8,150 884 Updated Jan 8, 2026

[ICML2026] Official Pytorch Implement for "Search or Accelerate: Confidence-Switched Position Beam Search for Diffusion Language Models"

Python 8 Updated Jul 6, 2026

Implementations of IQL, QMIX, VDN, COMA, QTRAN, MAVEN, CommNet, DyMA-CL, and G2ANet on SMAC, the decentralised micromanagement scenario of StarCraft II

Python 1,756 304 Updated Sep 8, 2022

Train transformer language models with reinforcement learning.

Python 19,040 2,901 Updated Aug 10, 2026

multi-agent deep reinforcement learning for networked system control.

Python 448 93 Updated Sep 29, 2020

Official implementation for "Unifying Masked Diffusion Models with Various Generation Orders and Beyond"

Python 6 Updated May 14, 2026

A framework for few-shot evaluation of language models.

Python 13,586 3,473 Updated Aug 10, 2026

Official Implementation for the paper "d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning"

Python 454 55 Updated Jan 26, 2026

Revisiting Discrete Gradient Estimation in MADDPG

Python 29 4 Updated Feb 24, 2023

dLLM: Simple Diffusion Language Modeling

Python 2,666 282 Updated Jul 17, 2026

Dream 7B, a large diffusion language model

Python 1,261 78 Updated Nov 21, 2025

Official PyTorch implementation for ICLR2025 paper "Scaling up Masked Diffusion Models on Text"

Python 385 28 Updated Dec 22, 2024

Multi-Agent Reinforcement Learning (MARL) papers

306 41 Updated Jul 10, 2026
Next