Skip to content
View jw9730's full-sized avatar

Highlights

  • Pro

Organizations

@kaistvllab

Block or report jw9730

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Understanding R1-Zero-Like Training: A Critical Perspective

Python 1,270 61 Updated Aug 27, 2025

Train transformer language models with reinforcement learning.

Python 19,052 2,903 Updated Aug 11, 2026

[ICLR 2026] SoFlow: Solution Flow Models for One-Step Generative Modeling

Python 163 7 Updated Apr 8, 2026

verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"

Python 2,206 212 Updated Jun 9, 2026

Implementation of Hindsight Differentiable Policy Optimization, as described in the paper Deep Reinforcement Learning for Inventory Networks: Toward Reliable Policy Optimization

Python 25 8 Updated Nov 19, 2025

Official codebase for the paper "How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance" (ICML 2026).

Python 15 Updated Jun 10, 2026

Parallel Token Prediction for Language Models (ICLR 2026)

Python 30 4 Updated Apr 27, 2026

Diffinity is a tool for constraining the output of continuous diffusion models to satisfy regular expressions. Companion artifact for ICML 2026 paper "Continous Diffusion Models can Obey Formal Syn…

Python 3 Updated May 25, 2026
Python 15 4 Updated Jul 14, 2026

Official Code Repo for Paper: Posterior Refinement

Python 10 1 Updated Jun 26, 2026

Official Implementation of MARS

Python 30 Updated Apr 21, 2026

Code for NeurIPS'24 paper 'Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization'

Python 240 26 Updated Jul 19, 2025

Code for the paper "Lessons from Studying Two-Hop Latent Reasoning"

Python 9 2 Updated Jun 18, 2026

A toy eval suite for tracing generalization dynamics of LM pre-training

Python 19 1 Updated May 19, 2026

Implementation of rewriting ensembles for the paper "What are the Right Symmetries for Formal Theorem Proving?"

Python 4 Updated May 27, 2026
Python 6 Updated May 17, 2026

Code of Training-free Detection of AI-generated images via Cropping Robustness

Python 8 1 Updated Jan 14, 2026
Python 32 2 Updated May 14, 2026

Official Pytorch Reimplementation of XFactor: True Self-Supervised Novel View Synthesis is Transferable (ICLR 2026, Oral)

Python 6 1 Updated May 17, 2026
Python 951 75 Updated Jun 26, 2026
Python 9 1 Updated Jun 3, 2026

Official implementation of "Infinite Mask Diffusion for Few-Step Distillation" (ICML 2026)

Python 6 Updated May 15, 2026

An LLM-agent framework that acts as a data scientist for relational learning.

Python 8 1 Updated May 12, 2026
Python 8 Updated May 14, 2026
Python 28 Updated Dec 19, 2025

Official implementation of Gumbel Distillation for Parallel Text Generation

Python 20 1 Updated Mar 24, 2026

Dreamer 4 jax implementation

Python 101 10 Updated Jul 24, 2026

A ~9M parameter LLM that talks like a small fish.

Python 3,331 291 Updated Apr 15, 2026
Next