LTX has launched LTX-2.5, an open weights world model optimized for local inference on NVIDIA RTX GPUs and DGX Spark, significantly reducing VRAM requirements for video generation, real-time applications, and physical AI. This release signals a shift towards local GPU processing over cloud infrastructure. Architecturally, LTX-2.5 incorporates native multishot generation to maintain visual consistency across sequences, integrates a Gemma 4 language backbone, and features a new decoder designed to minimize artifacts in high-motion footage. The model operates within ComfyUI and supports LoRA fine-tuning, providing direct control over IP and creative styles without external dependencies. The entire generation pipeline was rebuilt to achieve these capabilities.
In image-to-video generation benchmarks, LTX-2.5 demonstrates substantial performance advantages. A 10-second clip generates in 6.8 seconds on-prem using 2x NVIDIA GB200, and 23.7 seconds via the LTX API. This contrasts with closed alternatives such as Omni Flash, Grok 1.5, and Veo 3.1, which range from 52 to 70 seconds. Slower systems like Seedance 2.0, FLUX 3, Seedance 2.5, and Kling 3.0 Pro record times of 196, 259, 317, and 398 seconds, respectively. On-prem, LTX-2.5 is 7.6x faster than the nearest closed alternative and approximately 58x faster than the slowest, enabling practical overnight batch generation and rapid iteration. This aligns with NVIDIA's broader strategy to accelerate open models locally, positioning LTX-2.5 as a key component in a scalable, hardware-agnostic AI ecosystem, alongside releases like the Nemotron 3.5 Lightning 30B MoE agent model.
