MiniMax
MỖI GIÂY$0.0800/s
An omni-modal video generation model that can combine text with image, video, and audio references to create video assets with synchronized audio. It is suited to short-form video, advertising, and reference-driven character or camera work; it is not a general chat model, so prioritize visual consistency, motion, and audio-visual quality.
Kiểu đầu vào:
Kiểu đầu ra: