Skip to content
View fcjian's full-sized avatar

Block or report fcjian

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Open-source unified multimodal model

Python 6,116 544 Updated May 4, 2026

Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.

Jupyter Notebook 1,583 66 Updated Jun 14, 2025

NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.

Python 7,672 1,357 Updated Jul 22, 2026

[IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Python 3,106 212 Updated May 29, 2026

Janus-Series: Unified Multimodal Understanding and Generation Models

Python 17,754 2,232 Updated Feb 1, 2025

[NeurIPS 2025] Improving Video Generation with Human Feedback

Python 485 14 Updated Sep 24, 2025

A Next-Generation Training Engine Built for Ultra-Large MoE Models

Python 5,164 432 Updated Jul 23, 2026

NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.

Jupyter Notebook 11,226 788 Updated Jul 24, 2026

FastPillars: A Deployment-friendly Pillar-based 3D Detector

Python 181 13 Updated Jan 14, 2025
Python 31 Updated Jun 24, 2024

codes for RFSR: Improving ISR Diffusion Models via Reward Feedback Learning

Python 18 Updated Dec 8, 2024
Python 108 11 Updated Dec 27, 2024
Python 66 3 Updated Feb 20, 2025

A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

107 Updated Nov 28, 2024
Python 4,713 471 Updated Jun 15, 2026

OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

Python 408 33 Updated Mar 12, 2025

The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use th…

Jupyter Notebook 19,585 2,512 Updated May 30, 2026
8 Updated Jan 27, 2026
Python 630 51 Updated Mar 3, 2026

A curated list of awesome knowledge-driven autonomous driving (continually updated)

501 22 Updated Jun 7, 2024

UniMD: Towards Unifying Moment retrieval and temporal action Detection

Python 57 1 Updated Jul 5, 2024

[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.

Python 24,939 2,775 Updated Aug 12, 2024

A curated list of awesome LLM/VLM/VLA/World Model for Autonomous Driving(LLM4AD) resources (continually updated)

1,878 109 Updated Jun 22, 2026

InstaGen: Enhancing Object Detection by Training on Synthetic Dataset, CVPR2024

Jupyter Notebook 92 4 Updated Apr 9, 2024

✨✨Latest Advances on Multimodal Large Language Models

17,956 1,128 Updated Jul 2, 2026

Project Page for "LISA: Reasoning Segmentation via Large Language Model"

Python 2,667 208 Updated Feb 16, 2025

🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.

Python 21,444 2,399 Updated Jul 23, 2026

[ICLR 2023] Official implementation of the paper "DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection"

Python 2,825 310 Updated Jul 31, 2024

[CVPR 2023] SadTalker:Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

Python 13,968 2,667 Updated Jun 26, 2024
Next