-
KAIST
- South Korea
-
19:05
(UTC +09:00) - https://jw9730.github.io
- https://scholar.google.com/citations?user=kSJAiE4AAAAJ&hl=en
- in/jw9730
- @jw9730
Highlights
- Pro
Stars
Official codebase for the paper "How to Guide Your Flow: Few-Step Alignment via Flow Map Reward Guidance" (ICML 2026).
Parallel Token Prediction for Language Models (ICLR 2026)
Diffinity is a tool for constraining the output of continuous diffusion models to satisfy regular expressions. Companion artifact for ICML 2026 paper "Continous Diffusion Models can Obey Formal Syn…
Official Code Repo for Paper: Posterior Refinement
Code for NeurIPS'24 paper 'Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization'
Code for the paper "Lessons from Studying Two-Hop Latent Reasoning"
A toy eval suite for tracing generalization dynamics of LM pre-training
Implementation of rewriting ensembles for the paper "What are the Right Symmetries for Formal Theorem Proving?"
Code of Training-free Detection of AI-generated images via Cropping Robustness
Official Pytorch Reimplementation of XFactor: True Self-Supervised Novel View Synthesis is Transferable (ICLR 2026, Oral)
Official implementation of "Infinite Mask Diffusion for Few-Step Distillation" (ICML 2026)
An LLM-agent framework that acts as a data scientist for relational learning.
Official implementation of Gumbel Distillation for Parallel Text Generation
A ~9M parameter LLM that talks like a small fish.
A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
Synthetic pretraining data by rephrasing the web
The first continuous diffusion language model that rivals discrete counterparts on standard language modeling benchmarks like LM1B and OpenWebText.
The repository contains code for Adaptive Data Optimization