Skip to content
View xidulu's full-sized avatar
💔
Not in the mood of working
💔
Not in the mood of working
  • Johns Hopkins
  • Gotham City

Highlights

  • Pro

Block or report xidulu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Parallel-trained transformers with recurrent inference through shifted memory passes.

Python 2 Updated Aug 15, 2026

Open implementation of Attention Residuals (Kimi Team, arXiv:2603.15031)

Python 82 10 Updated Apr 30, 2026
Python 6 Updated May 20, 2026

Official Implementation for Pre-print "T2MLR: Transformer with Temporal Middle-Layer Recurrence"

Python 6 Updated Mar 16, 2026

Experiments on the impact of depth in transformers and SSMs.

Python 46 6 Updated Oct 23, 2025

Official repository for Amortized Factor Inference Networks (AFINs).

Python 5 1 Updated May 27, 2026

Curated collection of research on the limitations of next-token prediction and methods that go beyond it.

29 Updated Jul 10, 2026

Optimization benchmark for diffusion model training on dynamical systems

Python 10 Updated Nov 26, 2025

Fast Polar Decomposition for Muon

Python 173 14 Updated Jul 2, 2026

Code for the analyses in "The Supervision Horizon: An Exploration of the Mechanics of On-Policy Self-Distillation"

Python 4 Updated Mar 9, 2026
Dockerfile 1 Updated Apr 26, 2026

norm balancing optimizers

Python 3 Updated Feb 17, 2026

Code for "What really matters in matrix-whitening optimizers?"

Python 25 1 Updated Oct 31, 2025
Python 14 6 Updated Aug 17, 2026

Local coordinate frames

Python 1 Updated Mar 25, 2026
Python 1 1 Updated Jun 20, 2026

Code for the paper "Function-Space Learning Rates"

Python 23 2 Updated Jun 3, 2025

This notebook provides clean and minimal PyTorch implementations of classic pairwise preference models, including Bradley-Tery, Davidson, and Thurstone-Mosteller.

Jupyter Notebook 1 Updated Apr 22, 2025

[ICLR 2025] MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts

Python 279 12 Updated Oct 16, 2024

AI-powered SDK for Software Development

Python 56 12 Updated Apr 19, 2026

A framework for steering MoE models by detecting and controlling behavior-linked experts.

Python 36 7 Updated Sep 12, 2025

REAP: Router-weighted Expert Activation Pruning for SMoE compression

Python 478 83 Updated Apr 17, 2026
Python 29 10 Updated Apr 14, 2025

(ICLR'26 + Netflix) Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning

Python 53 7 Updated May 23, 2026

Qwen3 MoE Router Statistics | 큐웬3 MoE

Python 11 Updated May 16, 2025

[ICLR'26] A scalable, effective, and general way to augment LLMs with billion-scale knowledge graphs using very little GPU memory cost.

Python 25 4 Updated Jan 27, 2026

A family of open-sourced Mixture-of-Experts (MoE) Large Language Models

Python 1,694 86 Updated Mar 8, 2024

An extension of the nanoGPT repository for training small MOE models.

Python 284 33 Updated Mar 9, 2025
Next