-
university of science and technology of china
- https://orcid.org/0000-0003-1835-8468
Stars
Qwen-CUA: Native Computer Use for (Almost) Everything — a screenshot-driven agent that operates computers with keyboard and mouse, jointly developed by the Qwen Team and XLang Lab.
MacAgentBench: Benchmark agents where they actually work — on macOS.
让每一次引用都成为可解释的影响力 Turning Every Citation into Explainable Impact
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
Youtu-Tip: Tap for Intelligence, Keep on Device.
[Neurips 2025 Spotlight] Official repository for the paper: OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language Prompts
[CVPR 2025] DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
A stable & generalizable GRPO method for AR image generation
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
ResNeXt,ResNet, tensorflow,python
An implementation of Covariance Pooling, with the framwork of AlexNet and the dataset of UC Merced
All-day Semantic Segmentation & All-day CityScapes dataset
Use DenseNet40 for remote sensing image scene classification
Toneyaya / SimCIS
Forked from SooLab/SimCIS[CVPR2025] Rethinking Query-based Transformer for Continual Image Segmentation
Toneyaya / Part2Object
Forked from SooLab/Part2Object[ECCV 2024] The official PyTorch implementation of the "Part2Object: Hierarchical Unsupervised 3D Instance Segmentation".
Toneyaya / DDCOT
Forked from SooLab/DDCOT[NeurIPS 2023]DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models
[ICCV2023] CoTDet: Affordance Knowledge Prompting for Task Driven Object Detection
The official PyTorch implementation of the CVPR 2023 paper "Contrastive Grouping with Transformer for Referring Image Segmentation".
[ICCV 2025] HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets
[ICML 2025] This is the official repository of our paper "What If We Recaption Billions of Web Images with LLaMA-3 ?"
Repository for the BioCLIP 2 model project. [NeurIPS'25 Spotlight] BioCLIP 2 is a biological foundation model trained on TreeOfLife-200M. Despite the narrow training objective, BioCLIP 2 yields ext…
Official Implementation for paper "Unleashing the Power of Visual Foundation Models for Generalizable Semantic Segmentation"
[CVPR 2025 Highlight] SoMA: Singular Value Decomposed Minor Components Adaptation for Domain Generalizable Representation Learning
[CVPR22] Official Implementation of DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic Segmentation