-
University of Hong Kong (HKU)
- Hong Kong
-
03:17
(UTC +08:00) - https://wen-xin.info
- @_xwen_
Highlights
- Pro
Stars
Official Implemenation for RAEv2: Improved Baselines with Representation Autoencoders
Cambrian-P: Pose-Grounded Video Understanding
Claude code environment for laywers
Accompanying code for "Discovering State-of-the-art Reinforcement Algorithms" Nature publication
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
Project website for "Beyond Language Modeling: An Exploration of Multimodal Pretraining" paper.
The first multiplayer video world model in Minecraft
Sample LaTex file for HKU PhD thesis.
LaTeX Template for HKU MPhil and PhD Thesis
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
[CVPR 2026] Pixio: a capable vision encoder dedicated to dense prediction, simply by pixel reconstruction
Official code of "LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer"
Sudoku4LLM is a Sudoku dataset generator for training and evaluating reasoning in Large Language Models (LLMs). It offers customizable puzzles, difficulty levels, and 11 serialization formats to su…
PyTorch implementation of JiT https://arxiv.org/abs/2511.13720
Official implementation of "Continuous Autoregressive Language Models"
Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation
The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that sho…
A Walsh Hadamard Derived Linear Vector Symbolic Architecture 🔥
[NeurIPS'25] Official repository of Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
A comprehensive JAX/NNX library for diffusion and flow matching generative algorithms, featuring DiT (Diffusion Transformer) and its variants as the primary backbone with support for ImageNet train…
(NeurIPS 2025) Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation