A narrative music video created with a fully local generative AI production workflow.
The project was built around a simple challenge: creating not just a series of visually strong shots, but a coherent music video with consistent characters, visual language, environments, styling and camera direction across the entire sequence.
Character consistency was supported with custom LoRAs for both image generation and LTX 2.3, while the wider production challenge was maintaining continuity across framing, movement, lighting, wardrobe, performance and mood.
The video was produced locally on a single workstation, using:
LTX 2.3 - video generation
Z Image Turbo, FLUX.2 [klein], Qwen Image Edit - image generation and editing
ACE-Step 1.5 - music generation
ChatGPT 5.4 - lyrics and prompt development
What interested me most in this project was the shift from generating individual shots to directing a complete sequence. A strong standalone image or clip is relatively easy to judge in isolation; maintaining a consistent creative language across an entire piece is a very different production problem.
The project is part of my ongoing exploration of how generative AI can be integrated into real filmmaking workflows - not only for rapid visual experiments, but for complete narrative and commercial production.
Built with