Skip to content
View shiyongde's full-sized avatar

Block or report shiyongde

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Pioneering Automated GUI Interaction with Native Agents

Python 11,321 865 Updated Jan 27, 2026

[CVPR 2026 Highlight] WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning

Python 102 9 Updated Jul 29, 2026

Embedding model prioritized towards Multimodal RAG, overall + VisDoc double top1 on MMEB benchmark

Python 36 1 Updated Jun 16, 2026

The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that sho…

Python 11,273 1,703 Updated Jul 31, 2026
Python 9 Updated Aug 3, 2026
Python 53 8 Updated Oct 20, 2025

🔧Tool-Star: Empowering LLM-brained Multi-Tool Reasoner via Reinforcement Learning

Python 408 24 Updated Apr 3, 2026

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Foundation Models

Python 32 3 Updated Nov 7, 2025

基于多智能体LLM的中文金融交易框架 - TradingAgents中文增强版

Python 31,043 6,539 Updated Jul 24, 2026

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces

86 7 Updated Jun 6, 2025

Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.

Jupyter Notebook 666 59 Updated Jul 27, 2026

skip-vision: efficient and scalable acceleration of vision-language models via adaptive token skipping

Python 12 1 Updated Oct 31, 2025

Inference, Fine Tuning and many more recipes with Gemma family of models

Jupyter Notebook 304 48 Updated Apr 2, 2026

诺亚盘古大模型研发背后的真正的心酸与黑暗的故事。

11,541 1,300 Updated Jul 9, 2025

[NeurIPS 2025] 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding

Python 159 Updated Dec 9, 2025

[ICCV'25] Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

Python 70 1 Updated Jul 22, 2025

A new zero-shot framework to explore and search for the language descriptive targets in unknown environment based on Large Vision Language Model.

Python 80 6 Updated Nov 28, 2024

[IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Python 3,132 215 Updated May 29, 2026

[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型

Python 10,120 789 Updated Sep 22, 2025
Python 55 6 Updated Dec 23, 2024

NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.

Python 7,793 1,386 Updated Aug 10, 2026

Efficient Triton Kernels for LLM Training

Python 6,561 580 Updated Aug 11, 2026

A live stream development of RL tunning for LLM agents

Python 4,146 589 Updated May 5, 2026

Fully local web research and report writing assistant

Python 9,302 974 Updated Aug 5, 2026

🦉 OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

Python 20,080 2,300 Updated Aug 7, 2026

[NeurIPS2025 Spotlight 🔥 ] Official implementation of 🛸 "UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface"

Python 282 12 Updated Nov 5, 2025

(CVPR 2025 highlight✨) Official repository of paper "LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models"

Python 611 32 Updated Feb 4, 2026

FlashMLA: Efficient Multi-head Latent Attention Kernels

C++ 12,831 1,120 Updated Jul 28, 2026
Next