- Hong Kong
-
21:24
(UTC +08:00) - https://www.cse.cuhk.edu.hk/~zhpei23/
Stars
Official Implementation for the paper "d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning"
[ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.
WeDLM: The fastest diffusion language model with standard causal attention and native KV cache compatibility, delivering real speedups over vLLM-optimized baselines.
分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等
OpenSquilla — Token-Efficient AI Agent with same budget, higher intelligence density
Elevate your AI research writing, no more tedious polishing ✨
Academic Research Skills for Claude Code: research → write → review → revise → finalize
Open-source harness distillation: frontier teachers improve prompts, tools, validators, skills, and runtime policies around weaker models.
将博导十年科研经验炼化为可直接调用的 AI 技能。从 Idea 构思到论文投稿,你的 AI 科研副导师。
OpenCodeInterpreter is a suite of open-source code generation systems aimed at bridging the gap between large language models and sophisticated proprietary systems like the GPT-4 Code Interpreter. …
Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)
[ICLR'24 spotlight] An open platform for training, serving, and evaluating large language model for tool learning.
Train transformer language models with reinforcement learning.
Headless Slay the Spire 2 CLI — play the full game from a terminal.
DFlash: Block Diffusion for Flash Speculative Decoding
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
The implementation for the paper, FuseGPT: Learnable Layers Fusion of Generative Pre-trained Transformers.
[ACL 2026 Main] Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
[COLM 2026] Proactive Inference for Efficient Mixture-of-Experts
SCOPE: Self-evolving Context Optimization via Prompt Evolution - A framework for automatic prompt optimization
OpenClaw-RL: Train any agent simply by talking
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞