Skip to content
View kevin-ssy's full-sized avatar
  • Google DeepMind

Organizations

@torrvision

Block or report kevin-ssy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Open Frontier Intelligence

8,489 665 Updated Aug 6, 2026

MultiModal Audio Generation in Raw Waveform Space.

Python 153 10 Updated May 26, 2026

A toolbox for real-to-sim reconstruction and robotic simulation

Python 238 22 Updated Mar 31, 2026

Reference PyTorch implementation and models for DINOv3

Jupyter Notebook 11,196 930 Updated Jul 15, 2026

ViPE: Video Pose Engine for Geometric 3D Perception

Python 2,075 169 Updated Jun 9, 2026

[ICCV 2025] SpatialTrackerV2: 3D Point Tracking Made Easy

Python 993 54 Updated Feb 27, 2026

Code for ICCV'2025 (Best student paper honorable mention) "RayZer: A Self-supervised Large View Synthesis Model"

Python 446 20 Updated Nov 24, 2025

MAGI-1: Autoregressive Video Generation at Scale

Python 3,764 239 Updated Jun 17, 2026

State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!

Jupyter Notebook 2,346 160 Updated Apr 13, 2026

[CVPR 2025 Oral]Infinity ∞ : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

Python 1,587 93 Updated Apr 16, 2026

Simulation platform for general-purpose robotics & embodied AI learning.

Python 29,759 2,837 Updated Aug 17, 2026

Implement a ChatGPT-like LLM in PyTorch from scratch, step by step

Jupyter Notebook 102,849 15,759 Updated Aug 10, 2026

[NeurIPS 2024] Code release for "Segment Anything without Supervision"

Jupyter Notebook 503 27 Updated Nov 20, 2025

Autoregressive Model Beats Diffusion: 🦙 Llama for Scalable Image Generation

Python 1,965 95 Updated Aug 15, 2024

Official Implementation for CVPR 2024 paper: CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor

112 4 Updated Jun 23, 2024

Your image is almost there!

Python 7,612 435 Updated Jul 26, 2024

[CVPR2024, Highlight] Official code for DragDiffusion

Python 1,259 95 Updated Jan 29, 2024

[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". A…

Jupyter Notebook 8,728 571 Updated Nov 10, 2025

A family of lightweight multimodal models.

Python 1,052 75 Updated Nov 18, 2024

[RSS2024] A Multi-Modal Large Language Model with Retrieval-augmented In-context Learning capacity designed for generalisable and explainable end-to-end driving

Python 130 10 Updated Oct 7, 2024

Official Repo For OMG-LLaVA and OMG-Seg codebase [CVPR-24 and NeurIPS-24]

Python 1,352 55 Updated Oct 15, 2025

[ICLR 2024 Spotlight] Official implementation of ScaleCrafter for higher-resolution visual generation at inference time.

Python 507 28 Updated Mar 7, 2024

Official Code for MotionCtrl [SIGGRAPH 2024]

Python 1,499 83 Updated Feb 19, 2025

IMProv: Inpainting-based Multimodal Prompting for Computer Vision Tasks

Python 57 5 Updated Sep 26, 2024

[ICCV 2023 Oral] "FateZero: Fusing Attentions for Zero-shot Text-based Video Editing"

Jupyter Notebook 1,163 109 Updated Aug 14, 2023

[ICLR 2024] Real-Fake: Effective Training Data Synthesis Through Distribution Matching

Python 80 4 Updated Dec 9, 2023

[arXiv:2309.16669] Code release for "Training a Large Video Model on a Single Machine in a Day"

Python 138 12 Updated Aug 23, 2025

Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).

Python 7,265 509 Updated Oct 30, 2025
Next