Highlights
- Pro
Lists (1)
Sort Name ascending (A-Z)
Stars
香港入境处智能身份证预约配额监控看板:5分钟级检测+邮件/飞书放号通知(第三方工具,非官方)
Implementation of “Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?”
Hy3 (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency.
Hy3 preview (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
MiroEval: A benchmark and evaluation framework for deep research agents — 100 tasks (70 text, 30 multimodal) assessed across synthesis quality, factuality, and research process. 13 systems evaluated.
mini-claude-code: A minimal, readable Python re-implementation of Claude Code
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
OpenSeeker: A search agent with open-source data and models
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…
The official repo of "WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents"
Official Code Repository for the paper "Distilling LLM Agent into Small Models with Retrieval and Code Tools"
Agent-Omit: Training Efficient LLM Agents for Adaptive Thought and Observation Omission via Reinforcement Learning
[CVPR 2026🔥] Enhancing Spatial Understanding in Image Generation via Reward Modeling
[ICLR 2026] VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
We introduce VISTA-Bench, a systematic benchmark from multimodal perception, reasoning, to unimodal understanding domains. It evaluates visualized text understanding by contrasting pure-text and vi…
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning.
EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
Infinity-AILab / LongVT
Forked from EvolvingLMMs-Lab/LongVTIncentivizing "Thinking with Long Videos" via Native Tool Calling
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation.
[ICCV 2025] Hybrid Layout Control for Diffusion Transformer: Fewer Annotations, Superior Aesthetics.
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing [ICLR 2026]