Xuhui Zhou
AboutPublicationsCVBlogMore

Blog

Beyond Softmax: When Context Becomes a Learner

Jul 20, 2026

An interactive path from KV caches and linear attention to delta rules, fast weights, and test-time learning.

Policy Gradient From One Answer

Jul 17, 2026

A visual derivation of policy gradient for language models, walking from terminal reward to return, Q-value, and advantage.

Frontier Model Benchmark Matrix

Jul 16, 2026

A source-audited comparison of public benchmark reporting across six current frontier-model releases, including Kimi K3.

Thinking in RL

Apr 13, 2026

An opinionated tour through the algorithm tree of modern LLM RL — PPO, GRPO, REINFORCE, REINFORCE++, DPO, and the theoretical ideas that tie them together.

The Quest of User-Effective AI Agents

Nov 2, 2025

Exploring what makes AI agents truly effective for users, beyond benchmark performance.

The overlooked "bad" word list ☠️

Dec 15, 2024

Stop using outdated bad word lists. Use ToxicTrig instead for better toxic language analysis.