-
KAIST
- South Korea
-
09:25
(UTC +09:00) - https://jw9730.github.io
- https://scholar.google.com/citations?user=kSJAiE4AAAAJ&hl=en
- in/jw9730
- @jw9730
Highlights
- Pro
Stars
Understanding R1-Zero-Like Training: A Critical Perspective
Train transformer language models with reinforcement learning.
[ICLR 2026] SoFlow: Solution Flow Models for One-Step Generative Modeling
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
Implementation of Hindsight Differentiable Policy Optimization, as described in the paper Deep Reinforcement Learning for Inventory Networks: Toward Reliable Policy Optimization
Official codebase for the paper "How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance" (ICML 2026).
Parallel Token Prediction for Language Models (ICLR 2026)
Diffinity is a tool for constraining the output of continuous diffusion models to satisfy regular expressions. Companion artifact for ICML 2026 paper "Continous Diffusion Models can Obey Formal Syn…
Official Code Repo for Paper: Posterior Refinement
Code for NeurIPS'24 paper 'Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization'
Code for the paper "Lessons from Studying Two-Hop Latent Reasoning"
A toy eval suite for tracing generalization dynamics of LM pre-training
Implementation of rewriting ensembles for the paper "What are the Right Symmetries for Formal Theorem Proving?"
Code of Training-free Detection of AI-generated images via Cropping Robustness
Official Pytorch Reimplementation of XFactor: True Self-Supervised Novel View Synthesis is Transferable (ICLR 2026, Oral)
Official implementation of "Infinite Mask Diffusion for Few-Step Distillation" (ICML 2026)
An LLM-agent framework that acts as a data scientist for relational learning.
Official implementation of Gumbel Distillation for Parallel Text Generation
A ~9M parameter LLM that talks like a small fish.