All guides

AI Video line · stop 03 of 14 · 22 min · members

Image to video: what a still can become, and choosing a first frame that survives

What a still can and cannot become, and how to choose a first frame that survives the motion.

Free with an account

Sign in to read.

Membership is free: an account opens all 86 script pages. The Lab, Studio Canvas and the paid guides need the $99 pass, paid once. Already signed in on this browser? The page opens by itself.

01

The mechanism

The model extrapolates forward from one frame.

Everything it shows you that was not in that frame, it invented.

Image to video takes a still and continues it. It does not know what is behind the subject, what the far side of an object looks like, or what happens outside the frame — and where the motion requires any of that, it makes it up.

This single fact predicts almost every failure. Drift, morphing, objects that change shape, backgrounds that dissolve — all of them are the model inventing where it had no information.

So the craft is largely in choosing a source frame and a movement that keep the model inside what it can see.

02

Choosing the still

Clean, sharp, evenly lit, unambiguous.

The dramatic frame is usually the wrong source.

A source frame with deep shadow, motion blur or shallow focus contains ambiguity, and every ambiguous region is somewhere the model will invent during the move.

What survives well: sharp throughout, evenly lit, subject clearly separated from background, nothing cut off at the frame edge in a way that begs to be revealed.

What does not: heavy bokeh, blown highlights, complex overlapping objects, anything where you cannot tell what is in front of what.

If you are generating the still yourself, generate it for this purpose. A frame optimised as a still and a frame optimised as a video source are different pictures.

03

Choosing the movement

Movement that reveals nothing new is safest.

Ranked by how much the model has to invent.

  1. No movement, ambient only. Subtle life — cloth settling, light shifting. Almost always works.
  2. Slow push in. Reveals nothing outside the frame and no new surfaces. Very reliable.
  3. Slow pull out. Reveals the edges, which the model invents. Usually fine, occasionally strange.
  4. Pan or tilt. Reveals a whole side of the scene. Expect invention.
  5. Subject rotation or turning. Requires a surface never shown. Expect failure.

Design sequences from the top of that list. A piece built from pushes and ambient motion looks controlled; one built from rotations looks like a series of near misses.

04

Duration

Short clips hold. Long ones drift.

Error accumulates, so the end of a clip is always the weakest part.

Every frame is generated from the previous one's understanding, and small errors compound. The first second is usually excellent; the fifth may not be.

Generate shorter than you think you need and use the beginning. Where you need a longer shot, generate several and cut between them, or extend from a late frame as a new source rather than asking for one long run.

Always check the final frame before watching the clip. If it has drifted, the clip is unusable regardless of how good the opening was — and in motion you will not notice.

05

The prompt

Describe the motion, not the picture.

The image already says what is in frame.

A common mistake is writing a full image description as the prompt for image-to-video. The still already carries all of that, and re-describing it invites the model to reinterpret what it can see.

Prompt the change instead: what moves, how fast, and what the camera does. Plus the constraint block — what must stay identical.

The camera pushes in very slowly. The subject stays still.
The product keeps its exact shape, colour and label position.
No cuts. No additional objects enter the frame.

Short, about motion, with the preservation stated explicitly.

06

When it will not work

Some stills cannot be animated usefully.

Recognise them and change the plan rather than the prompt.

A still where the subject is cut off mid-object, where several elements overlap confusingly, or where the interesting content is in a heavily blurred region will not animate well no matter how it is prompted.

Neither will a frame that requires the subject to do something — a specific gesture, an expression change, an interaction. These need the movement to be generated from scratch, which is a different and much harder problem.

In both cases the answer is a different source frame, chosen or generated for the movement you want. That is a two-minute change and it is more effective than any number of prompt attempts.