-
Friedrich-Alexander-Universität Erlangen-Nürnberg
- Erlangen, Germany
Stars
[CVPR2026 Highlight] Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens https://arxiv.org/abs/2603.19232
DiffVP: Differential Visual Semantic Prompting for LLM-Based CT Report Generation
[CVPR 2025🔥] Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
🏛️ 三省六部制 · OpenClaw Multi-Agent Orchestration System — 9 specialized AI agents with real-time dashboard, model config, and full audit trails
Diffusion model for synthetic 3D CT scan video generation — IDEA Lab, FAU Erlangen-Nürnberg
Auto-regressive 3D CT volume generation using Latent Video Flow Matching with a Spatial-Temporal DiT (STDiT), conditioned on CT report embeddings.
Image-to-Image Translation in PyTorch
Code for NeurIPS 2024 paper - The GAN is dead; long live the GAN! A Modern Baseline GAN - by Huang et al.
You Only Denoise once or Average (YODA) - A diffusion-based 2.5D medical image translation model with noise-supression
3D U-Net model for volumetric semantic segmentation written in pytorch
Repo for MedSyn: Text-guided Anatomy-aware Synthesis of High-Fidelity 3D CT Images
A python package to streamline evaluation of unconditional image generation models
PyTorch re-implementation of FlowTok: Flowing Seamlessly Across Text and Image Tokens
This repo contains the code for 1D tokenizer and generator
Muon is an optimizer for hidden layers in neural networks
Open-Sora: Democratizing Efficient Video Production for All
Caption free adapter that maps DINOv3 image embeddings into CLIP space so you can do zero-shot text -> image or image -> text with CLIP’s text tower
[CVPR 2022] StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2
[ECCV 2024] Official PyTorch implementation of RoPE-ViT "Rotary Position Embedding for Vision Transformer"
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Learnable Fourier Features for Multi-Dimensional Spatial Positional Encoding
An implementation of 1D, 2D, and 3D positional encoding in Pytorch and TensorFlow