LeMiCa already supports accelerated inference for Z-Image and provides three optional acceleration paths that balance image quality and speed.
| Z-Image | LeMiCa-slow | LeMiCa-medium | LeMiCa-fast |
|---|---|---|---|
| 2.55 s | 2.19 s | 1.94 s | 1.78 s |
Note: The above numbers are example latency measurements for a single H800 GPU with a 1024×1024 resolution. Actual performance may vary depending on hardware and configuration.
Please refer to the original Z-Image project for base installation instructions.
# vanilla Z-Image (no caching / acceleration)
python inference_zimage.py
# LeMiCa acceleration modes
python inference_zimage.py --cache slow
python inference_zimage.py --cache medium
python inference_zimage.py --cache fast
# use an explicit numeric cache step
python inference_zimage.py --cache 8If you find LeMiCa useful in your research or applications, please consider giving us a star ⭐ and citing it by the following BibTeX entry:
@inproceedings{gao2025lemica,
title = {LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation},
author = {Huanlin Gao and Ping Chen and Fuyuan Shi and Chao Tan and Zhaoxiang Liu and Fang Zhao and Kai Wang and Shiguo Lian},
journal = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2025},
url = {https://arxiv.org/abs/2511.00090}
}We would like to thank the contributors to the Z-Image and Diffusers.