CT-Sora is a diffusion model for generating synthetic 3D CT scan video sequences. It is developed at the IDEA Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Germany.
The project adapts the Open-Sora framework and replaces the original video generation pipeline with a medical imaging focus: training on CT volumetric data to enable controllable, high-quality CT video synthesis.
- Backbone: MMDiT (Flux-style Multimodal Diffusion Transformer)
- VAE: HunyuanVideo causal 3D VAE for video encoding/decoding
- Training framework: ColossalAI (ZeRO, tensor parallel, sequence parallel)
- Experiment tracking: Weights & Biases (W&B)
| Stage | Script | Config | Description |
|---|---|---|---|
| Stage 1 | scripts/diffusion/train_huny.py |
configs/diffusion/train/stage1_new.py |
Text-conditioned CT video generation |
| Stage 2 | scripts/diffusion/train_huny_denoiser.py |
(denoiser config) | Video-to-video denoising refinement |
# Clone the repository
git clone https://github.com/WongJiayi/CT-Sora.git
cd CT-Sora
# Install dependencies
pip install -e .
pip install -r requirements.txt- Python >= 3.10
- PyTorch >= 2.4
- ColossalAI
- HuggingFace Transformers, Diffusers
- einops, decord, wandb
See requirements.txt for the full list.
Download the required pretrained weights and place them in the ckpt/ directory:
| Model | File | Source |
|---|---|---|
| HunyuanVideo VAE | ckpt/hunyuan_vae.safetensors |
Tencent HunyuanVideo |
CT-Sora expects video data referenced by a CSV file, where the first column contains paths to .mp4 video files:
/path/to/scan_001.mp4
/path/to/scan_002.mp4
...
To generate a CSV from a folder of videos, run:
python scripts/cnv/generate_csv.py --input_dir /path/to/videos --output /path/to/data.csvEdit configs/diffusion/train/stage1_new.py to set your data path and output directory, then run:
torchrun --nproc_per_node 8 \
scripts/diffusion/train_huny.py \
configs/diffusion/train/stage1_new.py \
--dataset.data-path /path/to/ct_videos.csvtorchrun --nproc_per_node 8 \
scripts/diffusion/train_huny_denoiser.py \
configs/diffusion/train/stage1_new.py \
--load /path/to/stage1/checkpointSet the load field in your config:
load = "/path/to/checkpoint/epoch-global_step"python scripts/diffusion/inference.py \
configs/diffusion/inference/256px.py \
--ckpt-path /path/to/checkpointCT-Sora/
├── opensora/ # Core library
│ ├── models/
│ │ ├── mmdit/ # MMDiT transformer backbone
│ │ ├── hunyuan_vae/ # HunyuanVideo causal 3D VAE
│ │ ├── dc_ae/ # DC-AE spatial autoencoder
│ │ └── vae/ # VAE utilities
│ ├── datasets/ # Data loading and sampling
│ ├── acceleration/ # Distributed training utilities
│ └── utils/ # Config, logging, optimizer, etc.
├── scripts/
│ ├── diffusion/
│ │ ├── train_huny.py # Stage 1 training
│ │ ├── train_huny_denoiser.py # Stage 2 denoising training
│ │ └── inference.py # Inference
│ ├── vae/ # VAE training and evaluation
│ └── cnv/ # Data conversion utilities
├── configs/
│ ├── diffusion/
│ │ ├── train/
│ │ │ ├── image_new.py # Base training config
│ │ │ └── stage1_new.py # Stage 1 config
│ │ └── inference/ # Inference configs
│ └── vae/ # VAE configs
├── assets/ # Demo images and example outputs
├── docs/ # Documentation
├── ckpt/ # Model weights (not tracked by git)
└── requirements.txt
Copyright 2025 IDEA Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Germany.
This project is a derivative work based on Open-Sora by HPC-AI Technology Inc., and is licensed under the Apache License 2.0. See LICENSE for the full text.
This project incorporates the following third-party components, each subject to its own license:
| Component | License | Source |
|---|---|---|
| Open-Sora | Apache 2.0 — Copyright 2024 HPC-AI Technology Inc. | Base framework |
| HunyuanVideo | Tencent Hunyuan Community License | VAE weights |
| FLUX | Apache 2.0 — Copyright 2024 Black Forest Labs | MMDiT architecture |
| EfficientViT | Apache 2.0 — Copyright 2023 Han Cai | DC-AE components |
| T5 | Apache 2.0 — Copyright 2019 Google | Text encoder |
| CLIP | MIT License — Copyright 2021 OpenAI | Text encoder |
Important — HunyuanVideo VAE: The VAE weights used in this project are derived from Tencent HunyuanVideo and are subject to the Tencent Hunyuan Community License Agreement. Usage of these weights must comply with that license, including geographic restrictions and acceptable use policies. See LICENSE for the full terms.
This work builds upon Open-Sora by HPC-AI Technology Inc. We thank the Open-Sora team for their open-source contribution to the video generation community.
If you use CT-Sora in your research, please cite:
@misc{ctsora2025,
title = {CT-Sora: Diffusion-based Synthetic 3D CT Video Generation},
author = {IDEA Lab, FAU Erlangen-Nürnberg},
year = {2025},
institution = {Friedrich-Alexander-Universität Erlangen-Nürnberg},
howpublished = {\url{https://github.com/WongJiayi/CT-Sora}}
}