-
Tsinghua University
- https://scholar.google.com/citations?user=S0Y2gJkAAAAJ
Highlights
- Pro
Stars
Executable, measurable, and reproducible AI4AI toward recursive self-improvement. Home of OpenMLE and Frontis-MA1.
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
[MICCAI26] Localization-Grounded Supervision: Revisiting Vanilla SFT of Large Vision-Language Models for Medical Image Analysis
A clinical-related and flexible radiology report evaluation framework.
A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io
UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms
RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography
MR-RATE: A Vision-Language Foundation Model and Dataset for Magnetic Resonance Imaging
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
[CVPR '26] CaptionQA: Is Your Caption as Useful as the Image Itself?
U-VLM: Hierarchical Vision Language Modeling for Report Generation
RadEval: A framework for radiology text evaluation
[ICLR 2026] Official implementation of "Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs"
Resource collection of medical agent for clinical dialogue and health
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use th…
The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
SEED-Voken: A Series of Powerful Visual Tokenizers
[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
[BIBM2025] Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model
paper list, dataset, and tools for radiology report generation
A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding