LTX-2 Model
A new standard for AI video generation
Delivered as an IC-LoRA on LTX-2.3. Generate directly in HDR or convert existing SDR footage to EXR. More grading latitude, more range, ready for real finishing pipelines.
Regenerate just one section — replace video, audio, or both — without re-rendering the whole clip.
Repurpose a video to any aspect ratio for social, filling new areas with generated content.
Generate video where voice, music, and sound effects define structure, pacing, and motion.Built for production-grade workflows that require precise, harmonious control over audio-led scenes - from podcasts and avatars to voice-driven clips -not one-off demos or talking heads.
Creative control that holds up under pressure. Structure, motion, camera behavior, and identity can be directed with intent rather than guessed by the model.
Depth-aware generation
OpenPose driven motion
Camera control
Stylistic and visual consistency
Audio to video
The model adapts to your worlds, characters, and creative DNA. Customization becomes part of the workflow, not a research project.
LoRA training support
Style LoRAs
Tools for upscaling, restoration, and detail recovery, powered by the model’s multi-scale rendering pipeline.
Detail upscaling
Recreate and generate elements of already existing videos. Edit with surgical precision.
Retake
Extend scene
Low-cost, high-speed generation for rapid iteration, storyboards, and previews.
Technical characteristics:
Higher detail and motion stability for final, commercial-grade renders.
Technical characteristics:
Research
Built on a distilled hybrid architecture, LTX-2 delivers significantly higher generation throughput without compromising visual fidelity. It outperforms smaller models like WAN 2.2 14B under identical settings, enabling faster iteration and high resolution video workflows on modern GPU hardware.
Asymmetric dual-stream DiT architecture. Joint audio-video generation with bidirectional cross-attention and modality-aware classifier-free guidance. Open weights and code.
Rebuilt VAE for sharper detail. 4x larger text connector for tighter prompt adherence. Native portrait, cleaner audio, HDR output, and precise control over motion and camera.
For academic teams pushing the boundaries of video generation, world simulation, and multimodal AI. Grants, model access, and research partnerships with the LTX team.