This is Ryuta Hamamoto from TIMEWELL.
With Sora 2 — an AI model that generates high-quality video from text — production work that once demanded specialist knowledge, expensive equipment, and a great deal of time has come within reach. But typing a prompt does not get you a finished piece. The more powerful the tool, the more the work shifts onto whoever is planning it.
This article follows how David Sheldrick, a DP and director, uses Sora 2 — a method he developed out of his own music video work. It applies the "shoot it all in one day" format he built up before COVID-19 directly to AI video generation: ideation, rendering, music selection, and the final edit. Read it as a piece about designing a process, not about operating a tool.
Opening up the creative: the first step in Sora 2 production
Production with Sora 2 begins by settling the creative foundation. This is not a loose brainstorm — it determines the quality and direction of the finished piece. Sheldrick argues for spending at least half a day, ideally a full day, exploring directions and experimenting before starting the assembly. What happens here sets the pace for everything downstream.
Using the Sora 2 Explore page for ideas and prompt study
For inspiration, Sheldrick points to Sora 2's "Explore" page, where videos generated and shared by users worldwide appear alongside the prompts that produced them.
The value is not just in browsing visual variety. It is a practical place to study how other people phrase things to achieve a particular look or mood. You can see, on real examples, how descriptions like "cinematic shot," "detailed skin texture," or "dynamic camera movement" change the output. Think of it as finding a style for your project and picking up the vocabulary to realize it, at the same time.
Defining the world: from period setting to visual elements
Running in parallel with idea collection is world-building. Before working out individual scenes (Location 1, Location 2…), you settle the atmosphere, style, and period that run through the whole piece.
Is it a historical period piece, futuristic science fiction, a pastoral landscape, or the inside of a palace? Once that is fixed, the direction for costume, art, props, lighting, and color follows. Sheldrick's example uses a clear theme — "18th-century Marie Antoinette" — and everything else develops out of that core.
Working with ChatGPT: expanding prompts into style presets
To turn the defined world into instructions Sora 2 can act on, Sheldrick recommends bringing in a large language model such as ChatGPT. You feed in the basic idea and ask something like: "Expand this prompt and make it more detailed for use in Sora 2 video rendering as a preset."
From the single phrase "18th century, Marie Antoinette," you get a prompt covering fabric (silk, lace, brocade), color palette (pastels with gold accents), lighting (soft natural light, the glitter of a chandelier), camera work (elegant dolly shots, close-ups), and mood (decadent, romantic, whimsical). That gets saved as a Sora 2 "style preset" and becomes the basis for holding a consistent look across the project.
Systematizing multiple creatives: the Marie Antoinette example
Once the overarching style is set, you define and organize the specific scenes and elements — what Sheldrick calls "creatives" — that play out inside that world. It is the same thinking as planning locations and shot types before a music video shoot, and it is the most important part of this workflow.
| Creative | Content |
|---|---|
| 1. Hair and makeup close-up | Enormous 18th-century wig, white powder makeup, tight on the model's expression |
| 2. Palace interior | Wide corridors and ballrooms; capturing architecture and interiors |
| 3. Hunting scene | An aristocratic pastime of the era — horses, costume, landscape |
| 4. Gardens | Manicured hedge mazes and geometric gardens, in the vein of Hampton Court Palace |
| 5. Horses | As an aristocratic motif: elegant movement, detail on the tack |
| 6. Kintsugi model | An original visual built on the Japanese aesthetic of repairing broken ceramics with gold |
Placing concrete elements under one large style umbrella lets you generate varied shots while holding the whole together. Each creative carries its own prompt, but the governing preset keeps them inside the same world. That structure is what makes rendering efficient and material easy to manage later.
Looking for AI training and consulting?
Learn about WARP training programs and consulting services in our materials.
Execution: rendering and music selection
With the direction set and the scene elements organized, you move to generating footage in Sora 2 — and, at the same stage, choosing the music that gives the piece its life.
Applying style presets and rendering efficiently
The detailed style prompt is saved through Sora 2's preset function. Paste it in from the "Manage Presets" menu and you no longer need to restate the style details every time you describe an individual scene.
What matters in rendering is running each creative many times. AI generation is probabilistic, so the same prompt yields slightly different output on every pass. Sheldrick applies the preset, enters a basic scene description (for example, "close-up of a Korean K-pop model getting her hair and makeup done"), and repeats. The result is a large pool of clips — some close to the intent, some unexpectedly good. It is a sensible collection strategy given how generation actually behaves. Presets are not fixed either; adjust them as you see results.
Adding dynamism: inserting dance sequence prompts
A distinctive move in this workflow is layering a "second prompt" about dance sequences on top of the basic scene prompt, to inject movement and energy.
Even for a hair and makeup scene, the description "wearing a huge 18th century Marie Antoinette wig, white powder makeup" gets paired with something like "bold camera shot of ethnically diverse K-pop couture fashion while dancing in unison, dancing in a Queen's bedroom, crunk dancing, street dance, dancing with attitude, dynamic dance, movement, dynamic music video camera work." Scenes that would otherwise sit still gain momentum. Naming specific dance genres, and mixing in abstract terms like "attitude" or "dynamic movement," pushes Sora 2 toward a wider range of interpretations.
Why the music matters
Music is not background here; it sets the atmosphere, the rhythm, and the emotional weight of the piece. Sheldrick considers AI music generation still immature at this point and works from high-quality stock music platforms instead (he uses Artlist.io).
Selection usually happens once enough footage has accumulated, or early in the edit. The chosen track goes on the timeline first, and the footage is then cut against its structure (intro, verse, chorus, bridge, outro), rhythm, and peaks. The music becomes the blueprint for the edit. Since cut points, scene lengths, and transition timing all end up following the track, choose it carefully.
Bringing it to life in the edit: from assembly to finish
With footage and music in hand, you reach the edit. What Sheldrick calls "assembly" is not simply lining up material — it demands judgment about synchronization, rhythm, and visual storytelling.
Building the timeline from the "sausage"
The first step is a technique he calls the "sausage": drop every generated clip onto the timeline in a single row. At this point you ignore music sync and cut timing; the goal is to see everything you have at once.
Next you group the material loosely according to the creative structure defined earlier — hair and makeup here, palace interiors there, gardens next to them. The creative structure now doubles as a guide to the edit's shape: what opens, what plays under the peak of the track. This early organization is what makes a large pool of AI-generated material workable.
Syncing to music: cutting on the beat
With the track on the timeline, you cut the footage to its rhythm and progression. Sheldrick pays particular attention to landing cuts on bass hits and drops.
Say you want the model to open her eyes exactly as the bass enters after a quiet intro. The generated clip will rarely be the right length, so you trim it and adjust the in-point precisely. Sora 2 output occasionally contains unintended cuts inside a clip, which means splitting it and removing the offending section. Listening closely to the beat, the melody, instrument fills, and lyrics — and setting cut points that answer them — is the core of the edit.
Speed adjustment and transitions
Beyond cutting, adjusting playback speed is useful. If you have a five-second clip for a three-second musical phrase, speeding it to roughly 167% fits it (Command + R in Final Cut Pro). Slow motion, conversely, emphasizes a movement or adds drama.
Speed is both a way to fit a duration and a way to control rhythm and energy. Fast cuts against slow motion produce a sequence with light and shade. Transitions work the same way: straight cuts plus the occasional fade or dissolve, used where they help the eye move.
How long assembly actually takes
This stage takes real time. By Sheldrick's account, going from the "sausage" to all clips placed against the music with a basic cut structure takes one to two hours. Completing the assembly took him around four hours in total.
The number shifts with the volume of footage, the complexity of the track, and the quality bar — but it is telling. Even when AI generates the material, turning it into something meaningful still costs human time and judgment. Which shots, in what order, cut where, held for how long: those choices decide how the piece reads.
Summary
Sheldrick's process is a working example of new technology meshing with the production discipline that came before it. Four things are doing the work:
- A clear vision and defined world — settle the style that runs through everything before touching individual scenes
- Structured creative development — concrete elements under one large style, giving you variety and coherence at once
- Close coordination with music — treat the track as the blueprint for the edit
- Iteration in the edit — generation got faster; assembly remains a human job
None of this is specific to music videos. Corporate promotional films, product introductions, branding content, short-form social video, training material — expression that was previously out of reach on time, budget, or expertise becomes a realistic option when Sora 2 is paired with a systematic workflow. Concentrating effort upstream, on concept and world-building, and using Sora 2 as a powerful visualization tool, is where the return looks best right now.
AI video generation is still developing, and moving quickly. Design your own purpose and process, then put the tool to work inside it. That is where this tends to land.
Reference: David Sheldrick's walkthrough — https://www.youtube.com/watch?v=0dhX84UkwFs
Working through AI adoption together
TIMEWELL offers an AI consulting service called WARP, delivered on a monthly cadence by people who have run DX and data strategy inside large enterprises.
The question this article raises — which parts of a process to standardize as your own — shows up in exactly the same form wherever AI meets real work, not just in video. We work through translating your own operations into something an AI can execute.
If you would rather start with a stocktake, the AI Readiness Check is a reasonable entry point — about three minutes. For a specific conversation, get in touch here.






