-
The Chinese University of Hong Kong, Shenzhen
- Shenzhen, China
-
08:58
(UTC -12:00) - https://1ke-ji.github.io/
Stars
MLEvolve is an open-source autonomous system for end-to-end machine learning algorithm design and optimization powered by progressive search and experience-driven memory.
Agentifying Patient Dynamics within LLMs through Interacting with Clinical World Model
[ICML'26] Scaling Long-Horizon LLM Agent via Context-Folding
MyPhoneBench: Do Phone-Use Agents Respect Your Privacy?
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
The lightweight framework for building agents
The largest open-source medical AI skills library for OpenClaw🦞.
Scalable toolkit for efficient model reinforcement
Code for the paper: Modular Retrieval for Generalization and Interpretation.
slime is an LLM post-training framework for RL Scaling.
An early research stage expert-parallel load balancer for MoE models based on linear programming.
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
MiniMax-M2, a model built for Max coding & agentic workflows.
The official repo of "WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents"
A simple yet powerful agent framework that delivers with open-source models
🏆 Top-1 on 5+ benchmarks | Web UI | Supports MiroThinker, Claude, Kimi, OpenAI
Writing AI Conference Papers: A Handbook for Beginners
A final sanity checklist to help your CS paper get accepted, not desk rejected.
🪐 🔧 Model Context Protocol (MCP) Server for Jupyter.
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
Kimi K2 is the large language model series developed by Moonshot AI team
A curated list of cutting-edge research papers and resources on Long Chain-of-Thought (CoT) Reasoning with Tools.
Reinforcing General Reasoning without Verifiers
Extrapolating RLVR to General Domains without Verifiers