An interactive path from KV caches and linear attention to delta rules, fast weights, and test-time learning.
A visual derivation of policy gradient for language models, walking from terminal reward to return, Q-value, and advantage.
A source-audited comparison of public benchmark reporting across six current frontier-model releases, including Kimi K3.
An opinionated tour through the algorithm tree of modern LLM RL — PPO, GRPO, REINFORCE, REINFORCE++, DPO, and the theoretical ideas that tie them together.
Exploring what makes AI agents truly effective for users, beyond benchmark performance.
Stop using outdated bad word lists. Use ToxicTrig instead for better toxic language analysis.