You finish the render. You play it back. The image is exactly right — the light, the motion, the texture. All of it. You’ve been staring at this thing for days, and it finally looks the way you had intended, or at least hoped for.
And yet, there is a hollowness to it, a lack of depth, an anchor to hold it in place. It’s missing… presence, atmosphere. Another playback, and then you figure it out, in a place you weren’t looking - because you were looking, and not listening.
The generated audio was there, but presence is hard to prompt, especially in AI video work. It’s difficult to convince that eye that the scene is real. But I am finding that it’s even harder to convince the ear that the scene is complete.
Not because the audio is wrong. Because the environment is incomplete, your ear knows immediately that something is off. It looks like a film, but it sounds like an afterthought.
And almost nobody in my AI film community is talking about sound, aside from Suno-generated music beds or pirated needle drops that fill the screen. A vibing soundtrack might tell you how to feel. But it rarely tells you where you are.
And where you are is everything. Presence is not a setting. It’s the hardest part of the work ahead. And the moment you start treating it that way, you stop being a prompt engineer and start being something closer to a director — one who makes work for an audience, not just for a render queue.
Direction is a decision your AI platform can’t make.
What’s circulating as AI video right now falls into two categories. There’s the genuinely sophisticated work — rare, usually unannounced, built by people with production instincts deep enough to know what they don’t yet have. And then there’s everything else. Beautiful clips. Technically impressive renders. Twenty seconds of motion that arrives, impresses briefly, and means nothing.
The difference between the two isn’t the model. It isn’t the platform. It isn’t even the prompt.
It’s direction.
Direction is the set of decisions that happen before, during, and after generation — decisions that most people making AI content aren’t making because they’ve confused the output for the work. They’ve generated something that looks finished and called it done. What they’ve actually made is a very expensive illustration. A still frame with motion applied. Something that has the grammar of film without the syntax of story.
I spent the better part of last week building a 3.5-minute opening video for an APAC VC firm’s AGM. Not a test render. Not a personal experiment. The first piece of communication a room full of investors and LPs would experience before the firm’s year began. The tone of that room for the day started with my three and a half minutes.
And at one point in the production workflow, I found myself standing on a stool in my home office with a recorder pointed at my air conditioning unit.
Sight and Sound.
What my video needed — specifically, in a segment where the visual language was minimal, interior, almost meditative — was not the sound the LLM had decided to generate with the click of a “Generate Audio” box. To be honest, I didn’t even try to generate other sound effects in AI tools like Suno or ElevenLabs. The sound I wanted, the atmosphere that the scene needed, lives in my office every day, that quiet, low hum of the air conditioner.
I recorded it. Forty seconds of near-silence that I would eventually cut under eight seconds of motion. And when I laid it in, the frame stopped being a render and became a place.
That is Foley work. That is the thing that professional sound designers do on actual film sets, recording footsteps and fabric, and the specific creak of a door hinge because the recorded version never sounds right. And it is, increasingly, the work that serious AI-augmented production demands — not because the tools are insufficient, but because storytelling assets are all around us, ready to be captured, tamed, and tasked, if you just pay attention.
What surprised me wasn’t that I needed to do it, but that nobody had told me I would. Every AI film tutorial I’ve encountered focuses on the image. The motion. The model. The pipeline. Sound is an afterthought, and spatial sound — the specific acoustics of a specific place at a specific time — is essentially invisible in the conversation. Which is exactly how you can always tell when you’re watching something built by someone who doesn’t actually know how film works. The eye forgives a lot. The ear forgives almost nothing.
One additional example – when I created a video for one of my clients, the story was told through the voice of a 9-year-old Singaporean girl. She was telling the story of finding a new home with her family. I had generated a variety of voices in ElevenLabs, which, to be honest, did a rather good job, but the voice still felt disconnected, not to the film, but to my ears. It was as if there was a distance between what I saw and what I heard.
My niece was visiting, and I asked her to record the VO for me. I ended up selecting her second take. It was full of life, of inflection, with natural breath and pauses that were imperfectly placed. Her real voice, paired with the AI-generated footage, closed that gap between sight and sound in a way I had not expected.
Direction Is Not a Prompt
In the previous piece I wrote about orchestration — about moving from prompt engineering to systems architecture — I said the system is only as intelligent as the thinking you build into it. This article is what that thinking looks like when you’re inside a production rather than designing the pipeline.
Direction, in any medium, is the accumulation of decisions small enough to seem trivial and important enough to determine everything. The angle of a light source. The texture of a surface in frame. Whether a camera move begins before or after the cut. Whether the sound enters a half-beat early or exactly on time. None of these is a dramatic choice. All of them are the difference between a scene and a frame.
In AI film production, those decisions sit in an interesting place. The model will make choices — it will decide, based on your instruction, how to light a face or move through a space. And it will often make those choices beautifully. What it will not do is understand why those choices matter to this story, in this moment, for this audience. That’s not a limitation of the technology. That’s the definition of direction. It requires intent. It requires a human being who knows what they’re trying to say and is disciplined enough to use every element of the frame to say it.
The people generating slop are most likely not making bad prompts. They’re making no decisions. There’s a meaningful difference.
What Patience Actually Costs
The other thing nobody tells you about serious AI video production is how long it takes.
Not because the tools are slow. Because the standard is moving. Because getting a camera move to feel intentional rather than procedural requires iteration that doesn’t have a shortcut. Because the difference between a cut that works and a cut that feels inevitable is something you feel before you can articulate it, and you keep going until it feels right. Whether you are building scenes from AI-generated keyframes, or building sound design from recorded source material, syncing it to image, adjusting it against motion — that is production work, and production work takes time.
I have watched the AI film conversation become increasingly obsessed with speed. With hacks. With the four-step tutorial that gets you from nothing to a clip in twenty minutes. And I understand the appeal. The tools genuinely are faster than anything we’ve had before. That speed is real, and it matters.
But speed is not craft. And the places where craft lives — in the decision about where the camera is before it moves, in the specific texture of the silence under a shot, in the choice to record forty seconds of air conditioning because the preset wasn’t true enough — those places don’t compress. They require what AI cannot provide: your attention. Your judgment. Your willingness to stand on a stool to capture a small sound that will eventually fill a room.
That’s not a romantic argument for practical production. It’s a structural argument for what separates work that lands from work that merely exists.
The Craft Didn’t Go Anywhere
In the orchestration piece, I argued that creative craft doesn’t disappear inside an AI system — it moves upstream, becomes structural, and governs everything downstream. In production, I’ve found it also moves sideways. Into disciplines you didn’t expect to need. Into Foley work and sound design. Into the precise grammar of camera intention. Into light decisions that have nothing to do with prompting and everything to do with what you understand about how light behaves in space and what it means when it behaves differently.
Your eye is not obsolete. Your understanding of pacing, of texture, of what a camera move communicates versus what it merely depicts — that is precisely the intelligence you are bringing to this work. The model generates the frame. You decide what it means. And the decision requires everything you know.
What I keep finding is that AI video production doesn’t lower the bar for craft. But it can hide it. The outputs can look polished enough that the absence of direction isn’t immediately obvious. Which is exactly why so much of what’s circulating feels vaguely unsatisfying in ways people can’t quite name. They’re watching frames without meaning. Motion without intention. Videos that have been rendered but not made.
The bar is there. It’s just invisible until you know what you’re listening for.
I work with creative leaders and production teams on AI-augmented storytelling — not the theory, but the practice. The pipeline decisions, the craft choices, the moments where the tools stop, and the direction begins. If that’s the conversation you need to have, let’s have it.


