Skip to content
View viakai's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report viakai

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

[CVPR Oral 2022] PyTorch Implementation for "Learning to Deblur using Light Field Generated and Real Defocused Images"

Python 129 8 Updated Dec 1, 2022

TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction

Python 349 23 Updated Jun 12, 2026
54 Updated May 6, 2026

Edit Banana: A framework for converting statistical formats into editable.

Python 5,453 360 Updated Aug 3, 2026

official github code for "SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing"

Python 151 5 Updated May 26, 2026

[CVPR2026] Exploring Spatial Intelligence from a Generative Perspective

Python 31 1 Updated Jun 3, 2026

Official Implementation of "AssemLM: Spatial Reasoning Multimodal Large Language Models for Robotic Assembly"

Python 15 1 Updated Jun 11, 2026

[ECCV2026] Anchor Forcing is a cache-centric framework for interactive streaming video generation that preserves visual quality and coherent motion across prompt switches

Python 27 Updated May 4, 2026

The code for Stroke3D | ICLR 2026

Python 35 2 Updated Feb 11, 2026

[ECCV 2026 Oral] DreamID-V: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer

Python 668 92 Updated May 22, 2026

[ACL'26] EvoToken-DLM (Beyond Hard Masks: Progressive Token Evolution for Diffusion Language)

Python 48 2 Updated Apr 7, 2026

code for aim-uofa.github.io

JavaScript 10 4 Updated Jun 16, 2026

This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).

320 17 Updated Feb 17, 2026

[ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Python 1,007 44 Updated Feb 10, 2026

Enjoy the magic of Diffusion models!

Python 12,886 1,264 Updated Aug 7, 2026

[CVPR 2025] A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

358 26 Updated Feb 27, 2025

[NeurIPS 2025 Spotlight] A Generalist Diffusion Model for Vision Perception

Python 321 14 Updated Sep 21, 2025

Frechet Video Distance metric implemented on PyTorch

Python 34 5 Updated Mar 22, 2020

Microsoft PowerToys is a collection of utilities that supercharge productivity and customization on Windows

C 137,595 8,469 Updated Aug 9, 2026
Python 20 Updated Mar 4, 2025

One-shot and Few-shot 3D Editing without Per-Scene Optimization

175 11 Updated Aug 21, 2025

A curated list of awesome papers for reconstructing 4D spatial intelligence from video. (arXiv 2507.21045)

515 29 Updated Jun 5, 2026

Official implementation of the paper "Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content".

Python 255 8 Updated Mar 19, 2025

[ICLR2026] Any-to-Bokeh is a novel one-step video bokeh framework that converts arbitrary input videos into temporally coherent, depth-aware bokeh effects.

Python 143 14 Updated Feb 4, 2026

Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).

Python 13,412 1,302 Updated Jun 26, 2026

[Arxiv] Discrete Diffusion in Large Language and Multimodal Models: A Survey

Python 387 4 Updated Apr 4, 2026

[3DV 2026] Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting

Python 161 8 Updated Dec 9, 2025

[ICML2026] ACTIVE-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO

83 1 Updated Apr 30, 2026

[NeurIPS 2025] Official Repo of Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Python 126 6 Updated Dec 3, 2025

Python client for Baidu Yun (Personal Cloud Storage) 百度云/百度网盘Python客户端

Python 8,571 1,410 Updated Apr 2, 2025
Next