Stars
Serenity-inspired Agent Skill for supply-chain bottleneck stock research
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
[ACMMM 2026] PLUME: Latent Reasoning Based Universal Multimodal Embedding
[Findings of ACL 2024]Optimal Transport Guided Correlation Assignment for Multimodal Entity Linking
The official implementation of COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence.
[ACMMM'25] Referring Expression Instance Retrieval and A Strong End-to-End Baseline
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…
✨✨Latest Advances on Multimodal Large Language Models
Interleaving Reasoning: Next-Generation Reasoning Systems for AGI
Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
Latest open-source "Thinking with images" (O3/O4-mini) papers, covering training-free, SFT-based, and RL-enhanced methods for "fine-grained visual understanding".
Reading notes about Multimodal Large Language Models, Large Language Models, and Diffusion Models
😼 优雅地使用基于 clash/mihomo 的代理环境
We introduce 'Thinking with Video', a new paradigm leveraging video generation for multimodal reasoning. Our VideoThinkBench shows that Sora-2 surpasses GPT5 by 10% on eyeballing puzzles and reache…
A Moral Evaluation Benchmark for Chinese Large Language Models
Python Implementation of Reinforcement Learning: An Introduction
AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs
A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for p…
免费大模型API,支持免费调用GPT、DeepSeek等主流模型,免费额度10000点,每日刷新!另付费价格最低官方1-2折!
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
Understanding R1-Zero-Like Training: A Critical Perspective