AI Video Prompts: How to Write Them in 2026
In short
A prompt for AI video is a written brief for the generator: who's in frame, what happens, how the camera moves, what the light looks like and how long it runs. In 2026 three models set the tone — Sora 2, Kling 3.0 and Google Veo 3.1 — and each has its own temperament. But they share the same base structure: subject → environment → action → light and style → camera move (always last). The key rules: one action and one camera move per shot, 2–3 short "beats" instead of ten, specifics instead of generalities. Below: a structural breakdown, a table of techniques, per-generator differences and the usual mistakes. Prices and access options are marked "as of July 2026", with no invented numbers.
What an AI video prompt is made of
A good video prompt reads like a brief for a film crew. Professional guides converge on this order of blocks:
- Subject — who or what is in frame, with 2–3 details of appearance.
- Environment — place, time of day, 3–5 details of the setting.
- Action — what the subject does: one verb, one movement.
- Light and palette — the light source and 3–5 anchor colours, so the shot doesn't drift in tone.
- Style — genre, optics, reference ("cinematic", "documentary", "35 mm").
- Camera move — placed at the very end of the sentence.
A beat is one short action inside the shot. The 2026 models hold 2–3 beats confidently; ask for five actions in a row and the picture falls apart. Duration and orientation (vertical or horizontal) are better set in the generator's settings than described in words.
How to describe camera movement
A camera move is how the "operator" travels relative to the scene. The models understand film vocabulary, but only one technique per shot: two moves in one prompt produce confusion.
| Technique | What the camera does | How to write it |
|---|---|---|
| Dolly in/out | Moves forward or back | "slow dolly in on the face" |
| Track/truck | Moves sideways, parallel to the subject | "camera tracks alongside the runner" |
| Pan | Horizontal rotation | "smooth pan left to right" |
| Tilt | Vertical tilt | "tilt up the skyscraper" |
| Crane/boom | Rises or descends | "crane rising above the crowd" |
| Static | Tripod, the frame doesn't move | "static shot, camera does not move" |
A separate word on "movement beats": describe what's visible — smoke drifting upward, flame bending in the wind, a runner leaning forward. Open-ended phrasing like "the camera moves somewhere" reads badly to the model.
How prompts differ for Sora, Kling and Veo
All three models eat the same structure but react differently. Broadly, per the guides available in July 2026:
| Parameter | Sora 2 | Kling 3.0 | Veo 3.1 |
|---|---|---|---|
| Strength | narrative, physics, sound | smooth motion, price | photorealism, 4K, colour |
| Prompt style | story-driven, "a scene in beats" | strict formula, camera last | precise cinematography terms, often JSON |
| Negative prompts | no | yes | limited/no |
| Native audio | yes | not on some versions | yes |
The practical takeaway: Sora likes a short script with a line of dialogue and sound, Kling likes the clean formula "subject → environment → action → light → camera", Veo likes an extremely precise description of optics and light, up to structured JSON for consistency across shots. A detailed capability comparison is in Sora vs Kling vs Veo, and an overview of the services themselves is in AI video generators 2026.
Table of techniques: what strengthens a prompt
| Technique | Why | Example phrasing |
|---|---|---|
| Subject anchor | keeps the face from drifting | "man, 40, grey beard, blue jacket" |
| 3–5 colour anchors | holds the shot's tone | "palette: amber, graphite, warm white" |
| A single light source | realistic shadow | "light from the window on the left, soft" |
| 2–3 beats | coherent action | "pours coffee, looks up, smiles" |
| Camera last | clean movement | "...slow dolly in" |
| Negative on 3–5 artifacts | fewer rejects | "no text, no distortions, no extra fingers" (where supported) |
A negative prompt is a list of what must not appear in frame. It works on Kling and Runway; blowing it up to 50 words is counterproductive — pick 3–5 artifacts you've actually seen in your own generations, not "everything at once".
Which mistakes ruin videos most often
- Too many actions. "Runs, jumps, falls, gets up" in one shot is guaranteed mush. Split it into separate scenes.
- No camera. With no movement specified, the shot comes out static and lifeless.
- Vague spatial words. "Somewhere nearby", "in the background" produce geometric distortions — be more specific.
- Two camera moves. One per shot, full stop.
- Trigger words. Innocent-looking words can trip moderation filters — rephrase.
- A negative-prompt dump. A long list of bans degrades the picture more than it helps.
The skill here is the same as with text prompts: structure, specifics, one idea per block. To build the base, start with the guide on how to write prompts and look through ready-made prompts for AI models.
What it costs and how to get access from Russia
Briefly and without invention, per open sources as of July 2026:
- Sora 2. The standalone Sora web version and app have been shut down since 26 April 2026; access remains inside ChatGPT (Plus and Pro plans) and through the API — per available sources, the API runs until September 2026.
- Kling 3.0. There's a free tier (around 66 credits per day) and paid plans from roughly $10 a month; the service doesn't accept Russian-issued cards directly.
- Veo 3.1. Premium quality (4K, native audio), priced from roughly $0.40 per second of generation.
Prices and terms change — check the official plans before paying and use only legally available access methods, without circumventing blocks. Access and payment from Russia get a separate breakdown in Kling AI: how to use it in Russia.
Where to start
Take one scene, write it out across the six blocks, set a single camera move and generate. Then change one parameter at a time — that reveals a model's temperament faster than another ten guides will. The same structured prompting that helps you steer code in Claude Code and Cursor through the Quest engine works for video too: a clear structure, one step at a time, a verifiable result. Voiceover and editing for finished clips are covered by AI text-to-speech tools and AI video editing tools.
FAQ
Should a prompt start with the camera or the subject?
With the subject and environment. Almost every model (Kling especially) reads camera movement better when it sits at the very end of the sentence. First who and where, then what they do, light and style — and only at the finish, how the camera moves.
Do you need a negative prompt?
Not always. Not every model supports one: it helps on Kling and Runway, while Veo's support is limited or absent. If you use one, keep the list short: 3–5 artifacts that genuinely spoiled your generations.
Why does my video come out static?
Most likely there's no movement in the prompt — neither camera nor subject. Add one camera technique (dolly in, pan) and one visible movement beat in frame: smoke, wind, a step. One explicit movement is usually enough.
Can one prompt produce a long scene with multiple shots?
As a rule, no: one prompt, one shot with 2–3 beats. Kling 3.0 has multishot, but it's more reliable to assemble a long video from separate generations and cut them together.
Is a prompt for an image the same as one for video?
No. Video adds two dimensions: time (the sequence of beats) and movement (the camera plus motion in frame). A good image prompt describes a moment; a video prompt also describes how that moment unfolds.