# AI Video Prompts That Hold Up Past the Third Second

URL: https://prexi.art/journal/ai-video-prompts
Type: blog
Locale: en
Published: 2026-10-05
Updated: 2026-10-05

---

> Stop describing pictures and direct the shot: five slots, real camera language and a lock line, plus three prompts to copy into Veo, Kling or Runway.

You typed "a cool cinematic shot of a city at night" and got eight seconds of melted neon soup. Welcome to the club. Good **ai video prompts** don't describe a picture. They direct a shot: who does what, where the camera goes, how long it lasts, what stays frozen. Below is the five-part structure we steal from every prompt that works, plus ready-to-copy examples you can drop into Veo, Kling or Runway today.

## Why your prompt melts after two seconds

A still image prompt describes one frame. A video prompt has to describe time. That's the whole difference, and it's the reason a Midjourney habit falls apart the minute you hit the video button.

When you only describe a look, the model has to invent the motion. It picks the laziest option: a slow zoom, a drifting face, background that quietly rearranges itself. You asked for a mood and got a screensaver.

The fix is boring and it works. Name the **subject**, one **action**, the **camera move**, the **light**, and one thing that must **not change**. Five slots. Miss one and the model fills it for you, usually badly.

Google's own [Veo prompting guide](https://cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide) lands on the same skeleton: subject, action, style, camera position and motion, composition, ambiance. We just collapse it so you can hold it in your head while you type.

## The five-slot formula, with a prompt to steal

Here is the frame. Write it in this order, in one or two plain sentences, and keep it under about 60 words for a first pass.

- 
**Subject**: one noun phrase with two specifics. "A courier in a yellow rain jacket", not "a person".

- 
**Action**: one verb, one beat. "Crosses the street and stops under a flickering sign."

- 
**Camera**: shot size plus movement. "Low angle, slow push-in."

- 
**Light and grade**: "sodium streetlight, wet asphalt, 35mm grain."

- 
**Lock**: what stays put. "Background signage stays static, no extra characters."

Steal this:

> A courier in a yellow rain jacket crosses an empty street at night and stops under a flickering sign. Low angle, slow push-in, 35mm film grain, sodium streetlight reflecting on wet asphalt. Signage and buildings stay static, no extra characters, no text.

Notice what is missing. No "epic", no "masterpiece", no "8K ultra detailed". Those words are leftovers from the image-model era and they add noise, not direction. Spend those tokens on the camera instead.

![Vintage camera on a tripod slider track in a dim studio, ready for a slow dolly move](https://fdzlnqpwsaniezitwiuw.supabase.co/storage/v1/object/public/cms-media/prexi/2026-10/d651e0-img1.webp)

## Camera language is the cheapest upgrade you have

If you only fix one slot, fix this one. Models have seen millions of clips labelled with film vocabulary, so real camera terms pull real behaviour. Vague ones pull the average.

A short cheat sheet that actually changes outputs:

- 
**Push-in / pull-back**: the safest moves. Start here.

- 
**Tracking shot**: camera follows the subject sideways or from behind. Great for walking scenes.

- 
**Orbit / arc**: powerful, also the move most likely to warp faces. Keep the clip short.

- 
**Static locked-off**: underrated. A tripod shot with only the subject moving is the most stable thing you can ask for.

- 
**Handheld**: adds life, adds jitter. Say "subtle" or you'll get earthquake cam.

One move per clip. Two simultaneous moves ("orbit while zooming and tilting up") is where physics goes to die. If you want a complex reveal, build it as separate shots and cut them together.

## Image-to-video prompts: describe the motion, not the picture

Starting from a still is the quiet cheat code of this whole craft. You lock the look with an image model first, then you animate. The video prompt stops being a description and becomes a stage direction.

The image already contains the subject, the light and the palette. Repeating them wastes words and sometimes fights the frame. Write only what **moves**:

> The woman turns her head slowly toward the window. Curtains lift in a light breeze. Camera holds still. Everything else in the frame stays exactly as it is.

That last sentence matters more than it looks. "Everything else stays as it is" is a cheap lock against the drifting backgrounds that give AI footage away.

If you're building references in Midjourney or Flux first, keep the composition simple: one clear subject, room around it for the camera to move. A cramped, busy still gives the video model nothing safe to push into.

![Blurred video editing timeline on a monitor above a notebook of camera angle sketches](https://fdzlnqpwsaniezitwiuw.supabase.co/storage/v1/object/public/cms-media/prexi/2026-10/f643a9-img2.webp)

## Prompts worth copying: three scenes, three jobs

Prompt libraries love volume. We'd rather give you three that each teach a different reflex. Remix freely, but start from these.

**The product beat.** Short, controlled, no humans to warp.

> Macro shot of a matte ceramic mug on a stone counter at dawn. Steam rises slowly. Slow pull-back reveals a sunlit kitchen. Soft window light, shallow depth of field. The mug shape and counter stay unchanged. No hands, no text.

**The mood piece.** One character, one gesture, lots of atmosphere.

> A lone figure with an umbrella walks away down a rainy neon alley at night. Locked-off wide shot, anamorphic flare, heavy rain, wet reflections, 35mm grain. The figure keeps walking at a steady pace. No cuts.

**The loop.** For backgrounds, hero sections, social posts that need to replay without a seam.

> Slow-drifting clouds over a mountain ridge at golden hour, static camera, gentle parallax in the foreground grass. Soft film grade. Motion is subtle and continuous, start and end frames look the same.

Each one follows the five slots. Each one ends with a lock. That's the pattern, and it's why they hold together past second three.

![Lone figure with an umbrella walking away down a rainy neon alley at night, cinematic frame](https://fdzlnqpwsaniezitwiuw.supabase.co/storage/v1/object/public/cms-media/prexi/2026-10/bfc6da-img3.webp)

## Sound, duration and aspect ratio: the settings people forget

Three details sit outside the five slots and still wreck clips when ignored.

**Duration.** Short clips hold together, long ones drift. If a model offers five to eight seconds, ask for one beat and stop. Need thirty seconds? Plan six shots, not one heroic generation. Editors have cut on action for a century for a reason.

**Aspect ratio.** Decide before you write, because composition depends on it. A slow push-in down a corridor reads beautifully in 16:9 and feels cramped in 9:16. For vertical, put the subject in the centre third and leave headroom for interface chrome. Say the ratio in the prompt if the tool lets you.

**Sound.** Some models now generate audio with the picture. If yours does, write the sound like you write the shot: "distant traffic hiss, rain on a metal awning, one muffled siren". Skip music requests unless you really need them. You'll want to score it yourself anyway.

Here's a version of the courier prompt with sound added:

> A courier in a yellow rain jacket crosses an empty street and stops under a flickering sign. Low angle, slow push-in, 35mm grain, 16:9. Ambient sound: steady rain, distant traffic, faint electric buzz from the sign. No dialogue, no text.

Same five slots, one extra line. That's the entire upgrade.

## Which model gets which prompt?

Same prompt, different model, different result. Don't expect one master prompt to travel. A few honest generalisations, no benchmarks pretended:

Veo handles long, descriptive, cinematic sentences well and can carry dialogue and sound cues when you write them in. Kling tends to reward clear physical action and camera instructions, and it's a comfortable home for image-to-video. Runway likes tight, restrained prompts and gives you more hands-on control once you're iterating.

If you'd rather not juggle three tabs, Higgsfield puts several video models behind one workspace. That's handy when you want to run the same prompt through two engines and compare, which is the fastest way to learn what each one actually listens to.

Our take: pick one model, learn its quirks for a week, then branch out. Hopping engines every time a clip disappoints teaches you nothing.

## Skip these habits (they cost you credits)

Some advice circulates because it sounds smart. It isn't. Skip:

- 
**Quality-word stacks.** "Ultra realistic, 8K, masterpiece, award-winning" does not make a better video. It dilutes the instruction.

- 
**Long, novel-length prompts.** Past roughly 100 words, models drop early details. Short and specific beats long and poetic.

- 
**Asking for text in the frame.** Signs, captions and logos usually warp mid-clip. Add titles in your editor.

- 
**Five actions in one clip.** Five seconds fits one beat. Split the rest into separate generations.

- 
**Rewriting from scratch every time.** Keep one working prompt, change one slot, compare. That's an actual experiment.

## How to iterate without burning your credits

Treat each generation as a test, not a lottery ticket. Change one slot at a time. If the motion is wrong, touch the action line. If the look drifts, tighten the lock. If the camera floats, swap "orbit" for "slow push-in".

Keep a running note of what worked, in whatever format you'll actually open again. Are.na boards work well for this, so does a plain text file. The point is that your third prompt should start from your second prompt's winner, not from zero.

And review clips muted first. Judge the motion on its own before you let a nice soundtrack forgive it.

## What to try tonight

Pick one scene from the three above, run it in a single model, and change only the camera line. Run it three times: static, push-in, tracking. Put the clips side by side. You'll learn more from that ten-minute experiment than from any list of a hundred prompts.

Copy the prompt, drop it in, and see what breaks. Then tell us which slot you'd fix first.

## FAQ

### What makes a good AI video prompt?

A good AI video prompt names one subject, one action, a camera move, the light and a lock line for what must stay unchanged. It describes time and motion, not just a single frame.

### How long should an AI video prompt be?

Aim for 40 to 80 words on a first pass. Past roughly 100 words, most models start dropping early details, so short and specific beats long and poetic.

### Do I need different prompts for Veo, Kling and Runway?

Yes, expect to adjust. Veo handles long cinematic sentences and sound cues, Kling rewards clear physical action, and Runway likes tight prompts. Keep the five slots and reword the rest.

### How do I write image-to-video prompts?

Skip describing the image, since it already holds subject, light and palette. Write only what moves, name the camera behaviour, and add that everything else stays exactly as it is.

### Can AI video prompts include sound and dialogue?

If your model generates audio, yes. Write sound like a shot: ambient texture, one or two effects, and dialogue in quotes if needed. Models without audio ignore these lines.

### Why do AI videos warp faces and backgrounds?

Usually the prompt asks for too much: several camera moves, many actions or crowded scenes. Use one move, one beat, a locked-off or slow camera, and a lock line for the background.