Constructing the Narrative
The creation of Cliff’s Edge: Shattered was not a linear path from A to B. It was a cycle. I treated the Large Language Model (LLM) not as a writer, but as a sounding board to help expand my initial, fragile ideas of isolation. The core challenge was maintaining a “Context Loop.” As shown in my workflow, the story influenced the lyrics, which influenced the song, which then circled back to dictate the visual pacing. This structure ensured that the final video didn’t just look good, but actually felt like the song it was accompanying.
The Translation Layer
Moving from a text storyboard to actual video is where most AI projects fall apart. I had to build a “translation layer” between my vision and the model’s capabilities. I realized early on that searching for the perfect generation is a trap. Instead, I focused on “Strategic Decision Making.” I generated options, picked the best potential candidate, and then adapted my direction based on what the AI gave me. It was a negotiation. I learned to accept happy accidents and work around technical limitations by treating them as creative constraints rather than failures.
The Human Assembly
While AI generated the raw materials, the final assembly was intensely human. The editing phase was where the “Architecture of Emotion” really happened. AI models do not understand narrative arcs or musical timing. That was my job. I spent days manually cutting clips, adjusting pacing to match the heartbeat of the track, and discarding beautiful shots simply because they didn’t serve the story. The tools provided the bricks, but I had to lay them by hand to ensure the final structure stood up emotionally.
To explore the complete technical breakdown and read the full production diary, you can read on civitai article.












