Stars
AI-Generated Presets for Faithful 4K Color Style Transfer in Real Time [CVPR 2023]
Official Implementation of "Control-A-Video: Controllable Text-to-Video Generation with Diffusion Models"
This repository provides motion datasets collected by Bandai Namco Research Inc
[ICCV 2023] Few shot font generation via transferring similarity guided global and quantization local styles
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models (ICLR 2024)
Official PyTorch implementation of "EdgeSAM: Prompt-In-the-Loop Distillation for On-Device Deployment of SAM"
[CVPR2024] StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On
[NeurIPS 2024] SlimSAM: 0.1% Data Makes Segment Anything Slim
[CVPR 2024] Alpha-CLIP: A CLIP Model Focusing on Wherever You Want
[ICLR 2024] Official implementation of DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior
[ICLR'24 Spotlight] Uni3D: 3D Visual Representation from BAAI
GeoDream: Disentangling 2D and Geometric Priors for High-Fidelity and Consistent 3D Generation
Instant voice cloning by MIT and MyShell. Audio foundation model.
[CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
[ECCV 2024] FreeInit: Bridging Initialization Gap in Video Diffusion Models
Let us democratise high-resolution generation! (CVPR 2024)
[ECCV 2024] Official repo for UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion Models
ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation (NeurIPS 2023 Spotlight)
A unified framework for 3D content generation.
Refine high-quality datasets and visual AI models
End-2-end speech synthesis with recurrent neural networks
[ICASSP 2024] ๐ต Matcha-TTS: A fast TTS architecture with conditional flow matching
๐ธ๐ฌ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
serp-ai / bark-with-voice-clone
Forked from suno-ai/bark๐ Text-prompted Generative Audio Model - With the ability to clone voices
[ACL 2024] An Easy-to-use Knowledge Editing Framework for LLMs.
Official implementation of paper "MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens"
ModelScope: bring the notion of Model-as-a-Service to life.