AI Prompts for Video: Sora, Runway and Veo
Text-to-video models reward the language of film, not tags. Here's how to prompt Sora, Runway and Veo for shots that actually look intentional.

Type "a dragon flying over a city" into an image model and you get one frame. Type it into Sora or Veo and you get five seconds of decisions: how the dragon banks, whether the camera chases it or holds still, how the wings push air, where the light shifts as it turns. Every one of those decisions is either something you specified or something the model guessed for you. Good prompting for video is mostly the work of guessing less.
That changes how you write. A still is a noun. A shot is a verb.
Video Prompts Are Not Image Prompts With Motion
The biggest mistake people carry over from Midjourney is treating a clip like a photo that happens to wiggle. It isn't. A photo prompt describes a state. A video prompt describes a state *and how it changes over the duration*. You're now responsible for three things a still never asked of you: motion (what moves and how), time (what happens first, then next), and pace (how fast any of it unfolds).
This is why AI prompts for video live or die on verbs. "A woman in a red coat, neon street, rain" is a fine image brief and a weak video brief, because nothing in it moves. Add the motion and it wakes up: "she walks toward camera, coat swaying, rain streaking past neon." Same scene, but now the model knows what to animate instead of inventing drift on its own.
Keep the same discipline you'd use for any strong prompt — the fundamentals of writing a clear instruction still apply. You've just added a time axis, and the time axis is where clips go wrong.
The Anatomy of a Shot Prompt
A shot prompt has parts, and naming them in a rough order keeps you from forgetting the ones models fumble. Think of it as a slate:
- Subject + action — who or what, doing exactly one thing.
- Camera angle and movement — where the lens sits and how it travels.
- Lens and framing — wide, 35mm, close-up, shallow depth of field.
- Lighting — direction, quality, colour temperature.
- Mood and setting — the place and the feeling.
- Duration and pace — roughly how long, and how fast the action reads.
The best AI prompts for video rarely fill all six slots — but they always nail the first two, because subject and camera are load-bearing. Here's the shape filled in:
Notice there's one subject, one camera move, one time of day. That restraint is the whole trick. If you're used to stacking descriptors for stills, an image prompt builder trains the habit of listing lens and light explicitly — carry that over, then add the movement clause the still never needed.
Speaking the Camera's Language
Video models were trained on real footage, so they respond to real film vocabulary far better than to vague words like "dynamic." Learn a handful of camera moves and your text to video prompts get sharper overnight:
- Dolly — the camera itself moves forward or back ("slow dolly-in").
- Pan / tilt — the camera pivots left-right or up-down from a fixed spot.
- Tracking — the camera travels alongside a moving subject.
- Crane / jib — the camera rises or descends through space.
- Handheld — deliberate shake for documentary energy.
- Zoom — the *lens* tightens, which looks different from a dolly and models know it.
Name the move, name its speed, and pick one. "Slow tracking shot" reads clean; "dolly-zoom-pan-crane" reads like noise and the model will average it into mush.
Describing Motion and Physics
The hardest thing for these tools is physical plausibility — cloth, hair, water, weight, momentum. You can steer it. Instead of "a cape," write "a heavy wool cape lifting slowly in the wind." Instead of "explosion," write "a slow-motion burst of embers drifting upward." Give the model direction, speed, and material and it has something to simulate. Leave those blank and it defaults to a soupy, floaty average.
A few levers that pay off:
- Speed words: "slow-motion," "real-time," "time-lapse," "gradual," "sudden."
- Weight cues: "heavy," "delicate," "sluggish," "snapping back."
- Direction: "drifting left to right," "falling toward camera," "rising."
Contradictions are what break it. "A frozen still shot with fast chaotic motion" hands the model two incompatible goals, and it picks one at random.
One Shot, One Idea, Then Cut
You cannot brief a whole scene in a single prompt and expect coverage. Ask for "a chase through a market, then a rooftop standoff, then rain" and you'll get a confused blur that's none of them. Real edits are built one shot at a time, so prompt one shot at a time.
For a sequence, storyboard in plain language first — three or four beats, each its own generated clip — then assemble them in an editor:
- Shot 1: wide establishing, subject small in frame.
- Shot 2: medium, subject reacts.
- Shot 3: close-up, the detail that matters.
Consistency across those shots is the real difficulty. Repeat the anchoring description verbatim — same wardrobe, same location, same light — changing only the framing and action. A saved prompt library earns its keep here, because you'll reuse that character and setting block across every clip and small wording drift causes visible continuity breaks. The same logic that keeps an illustration set looking like siblings keeps your shots looking like one scene.
Image-to-Video and Reference Frames
Text isn't your only input. Most of these tools accept a starting image and animate it, which solves consistency in one move: you lock the look as a still, then describe only the motion. This is the most reliable way to control character and style, because the frame carries everything words struggle to pin down.
The workflow is worth building around. Generate a strong first frame in an image model — the image-prompt guide and the Midjourney prompt generator cover getting that frame right — then feed it in with a motion-only instruction:
Because the frame is fixed, the prompt only has to answer "what moves?" — a far easier question than "what does everything look like?"
Sora, Runway, Veo, Kling: Reading the Differences
The tools overlap more than the marketing suggests, and versions change fast, so treat specifics as moving targets. Broadly, though, they have different temperaments:
- Sora (OpenAI) leans toward longer, more narrative clips and reads scene descriptions well. Strong Sora prompts tend to be written like a shot list, not a keyword salad.
- Runway grew out of an editing toolset, so it pairs generation with fine motion controls, camera settings, and image-to-video.
- Google Veo aims at high-fidelity, physically plausible motion and responds well to explicit cinematography language.
- Kling and Luma are strong on realistic movement and are common go-to tools for image-to-video.
Don't over-fit to one tool's exact syntax; the underlying craft — clear subject, one camera move, named motion, sane duration — transfers everywhere. If a prompt reads awkwardly, run it through a prompt optimizer to tighten the wording before you spend another render.
Mistakes That Burn Renders
Clips cost time and credits, so the failures worth naming are the ones that waste both:
- Overloading. Six actions, four camera moves, three moods. Cut to one of each.
- Contradictory directions. "Static locked-off shot, wild camera movement." Pick a lane.
- Vague "cinematic." The word does almost nothing alone. Say *what's* cinematic — the anamorphic flare, the low-key lighting, the slow push. Specificity is the style.
- Ignoring duration. A five-second clip can't hold a three-beat story. Match the ambition to the length.
- Fighting the still layer. If the first frame looks wrong, the motion won't save it — fix the frame first.
Write like a director filling a slate, not a poet describing a dream. The models are good listeners and terrible mind-readers, and the best AI prompts for video simply leave less to guess.


