Animate a photo into a video clip
The photo is the hard part — the face, the look, the scene. Animation takes that still as the literal first frame of a clip and adds only the motion: a turn of the head, hair catching wind, the crowd shifting behind her.
Your still is the first frame
Image-to-video doesn't invent a new scene — it takes your exact photo and moves it. Whoever is in frame stays who they are; the model's job is confined to what changes over time.
Attach the still, write the motion, pick duration and resolution, run it. A render takes a few minutes — it's the slowest job type in the studio, and billing is per second of output.
Clips run 4–30 seconds depending on the model, at up to 4K on the newest tier, with an audio track you can toggle off.
Prompts that work
"She turns to the camera and smiles subtly, hair moving in the breeze." — one subject action plus one ambient motion. That pairing is the reliable shape.
"Slow push-in, the crowd behind her shifts, she holds the pose." — camera moves are valid instructions; the subject can stay nearly still while the frame breathes.
"Confetti falls, she laughs." — ambient particles plus a subject reaction reads naturally.
"Waves keep rolling, lighthouse beam rotates slowly, clouds drift." — for scenes with no person, list two or three independent motions.
Write motion, not a picture
The most common failed prompt describes a scene instead of a movement — "a woman at a beach bar at sunset" is the first frame, which the model already has. Write what happens next: "she raises her glass", "the lights behind her flicker on".
One clear motion beats three competing ones. "Turn + smile + walk away" tears the choreography; "she turns toward camera" lands. Add the next motion as a second clip if you need it.
Camera language works: push-in, pull-back, slow pan, orbit. Pair at most one camera move with one subject action.
Spending tips
Render the first pass at 480p and 4–5 seconds — it costs a fraction of the flagship settings and tells you whether the motion reads right. Only re-render the winning prompt at higher resolution and length.
Kill the audio if you don't need it — on the models that charge for the track it's a real saving, and you can always re-run with sound once the motion is locked.
Iterate the prompt cheaply upstream: get the still right with image edits first (a few cents each), because every i2v re-render costs more than fixing the source frame.
The honest limits
Motion can add a turn or a breeze, but it can't walk a seated figure across a room — ask for movement the pose can plausibly start.
Longer clips drift more: the last seconds of a 30s render are the least faithful to the first frame.
Clips land in the same local library, auto-deleted after 48 hours — download what you keep.