🌟 Align diffusion processes with detailed human preferences to improve machine learning models for richer, more accurate outputs.
-
Updated
Sep 23, 2026 - Python
🌟 Align diffusion processes with detailed human preferences to improve machine learning models for richer, more accurate outputs.
A collection of robotics simulation environments for reinforcement learning
Split conformal prediction for off-policy evaluation in offline RL (D4RL MuJoCo). Stage 1 of 4: episode reconstruction and data foundation, with guards for two silent D4RL failure modes.
Robot RL portfolio: bipedal locomotion, algorithm comparison, offline RL, gait phase estimation (4 reproducible projects)
Diagnostic study of regime representations in offline reinforcement learning under heterogeneous behavior data.
[IcETRAN 2026] Official implementation of Flow Matching Policy for Behavioral Cloning paper.
Clean single-file implementation of offline RL algorithms in JAX
ANA 699 Capstone — Offline reinforcement learning for robotics using Decision Transformers and MuJoCo
Unified offline RL / offline-to-online RL course project scaffold with D4RL smoke tests
Offline RL via sequence modeling: BC, Decision Transformer, and Online DT comparison on D4RL benchmarks
Online goal-reaching RL with diffusion planning: a 2D point-mass prototype plus a Maze2D Diffuser workflow driven by an autonomous agentic experiment controller.
D4RL benchmark but ported to work end to end with gymnasium
[NeurIPS 2025] A human-like RL framework that improves human-likeness while achieving strong performance, and can be easily integrated into various RL algorithms
Unified Implementations of Offline Reinforcement Learning Algorithms
Diffusion-Guided Tree Search: Uncertainty-Aware Planning with Learned World Models
Velocity-parameterized diffusion for trajectory planning, with MPPI sampling and periodic replanning. Matches or beats Diffuser (ICML 2022) on Maze2D at half the diffusion steps.
a clear and fast jax/flax version of [Diffusion-Policies-for-Offline-RL](https://github.com/Zhendong-Wang/Diffusion-Policies-for-Offline-RL)
Codes accompanying the paper "Score Regularized Policy Optimization through Diffusion Behavior" (ICLR 2024).
Learning from Sparse Offline Datasets via Conservative Density Estimation (ICLR 2024)
High-quality single-file implementations of SOTA Offline and Offline-to-Online RL algorithms: AWAC, BC, CQL, DT, EDAC, IQL, SAC-N, TD3+BC, LB-SAC, SPOT, Cal-QL, ReBRAC
To associate your repository with the d4rl topic, visit your repo's landing page and select "manage topics."