Prompt Guide
Image to Video Prompts That Preserve Your Original Image
Use a reliable image to video prompt structure for controlled camera motion, natural movement, and fewer unwanted changes.
The best image to video prompt describes change over time. The source image already defines the subject, setting, color, and composition, so focus on motion, timing, camera behavior, and the elements that must remain stable.
Use a motion first prompt structure
Start with the camera instruction, then the main action, environmental motion, timing, and preservation constraints.
Example: Locked camera. The character turns toward the window and takes one slow step. Curtains move gently. Preserve the face, outfit, room layout, and lighting. No new objects.
Choose camera movement deliberately
A locked camera is the safest start for identity and product accuracy. A slow push in adds energy without revealing much new scene information. Orbits and large pans require the model to invent hidden areas.
- Locked camera for dialogue and precise blocking
- Slow push in for emphasis
- Pull back for a controlled reveal
- Pan only with enough surrounding space
- Orbit only when invented surfaces are acceptable
Use first and last frames for exact destinations
When a shot needs a specific final pose or composition, create both endpoints. The first establishes the starting state and the last defines where motion finishes.
Keep the frames compatible. If camera angle, identity, lighting, and background all change, divide the transition into shorter shots.
Diagnose common failures
If the face changes, reduce camera movement and action. If the background melts, lock the camera and remove unnecessary environmental motion. If nothing moves, use one clear physical verb.
Negative instructions are guardrails, but they cannot rescue a contradictory prompt. Simplify desired motion first.
Frequently asked questions
What should an image to video prompt include?
Include camera behavior, subject action, environmental motion, timing, and a short list of details that must remain unchanged.
Why does image to video change the original face?
Large movement, camera rotation, weak references, and competing details can force the model to reconstruct the face.
When should I use first and last frame video?
Use it when a shot must end in a specific pose, composition, location, or product state.