Eight qualified image models. One prompt-to-download workspace.

PRACTICAL GUIDE · PROMPTTOVISUAL

AI Video Keyframe Control: First, Middle, Last Frames and Camera Paths

Image keyframes constrain what a generated video should show at chosen moments. Reference images guide appearance, while camera keyframes describe a viewpoint path. Choose the control that matches your task before preparing inputs.

Choose controls for the task

Task selection — documented controls, not tested output guarantees
Your taskControlWhat to prepare
Keep a subject or look recognizableReference imagesIdentity or style material; not automatically timed anchors
Animate an existing stillFirst frameOne starting picture and a modest motion instruction
Connect two specific compositionsFirst and last framesA plausible start and destination
Include a particular intermediate compositionFirst, middle and last framesThree compatible pictures plus an interior time; inspect the H3 Max reference-to-video contract
Move the viewpoint around a sceneCamera trajectory keyframesCamera poses along a normalized timeline, not three uploaded images

First and last frames: choose start and destination

For H3 Max on fal, minimax/h3-max/image-to-video documents image_url for the first frame and end_image_url for the last frame. This is a useful contract to inspect when your task has two visual endpoints. A starting picture alone leaves the destination to the model and prompt; two endpoints still do not specify every intervening action.

Source for this section →

First, middle and last: use the correct endpoint

The three-image contract is minimax/h3-max/reference-to-video. Use image_url, middle_image_url and end_image_url; middle_frame_time sets the intermediate time in seconds. A middle image requires both start and end images and native 480P or 768P output. Its time is rounded to the nearest frame at 24 fps and must fall strictly between the first and last frames of the requested duration. These are documented constraints, not a workflow we have executed.

Source for this section →

Camera keyframes are a different input

minimax/h3-max/camera-controls takes an initial image_url and a camera_trajectory of poses. Each keyframe has time normalized from 0 to 1, azimuth and elevation in degrees, and distance in normalized scene units. Do not send middle_frame_time seconds as normalized camera time or treat three pose entries as three uploaded pictures.

Source for this section →

Prepare inputs before paying for an attempt

Use images with a consistent subject, proportions and framing. Check that the intended movement can plausibly connect them within the chosen duration. Keep image preparation separate from motion instructions: label which picture is the start, the intermediate target and the destination. For reference-only inputs, describe the identity or style to preserve rather than assuming the upload order sets a timeline. These are preparation recommendations, not guarantees of exact preservation.

An original timing example — untested planning exercise

Imagine a five-second clip of a toy car crossing a tabletop. Plan a starting image at the left edge, an intermediate image near the center at 2.5 seconds, and a final image near the right edge. The intermediate time is inside the clip, and the pictures describe a plausible sequence. This is an untested planning example, not generated evidence, a ready-to-run request or a guarantee that the car follows an exact path. Review the endpoint’s current duration and resolution settings before using it.

Common reasons a keyframe plan fails

A reference upload may influence appearance without fixing a moment. Incompatible poses may invite a visible morph instead of the intended motion. Too much action in a short interval can defeat a clear transition. A request can also fail validation when it uses the wrong endpoint, omits required companion images or mixes time units. Separate input validation from output quality: a successful request still needs visual inspection.

Where should I go next?

For three-frame control, read the external H3 Max reference-to-video contract linked below. PromptToVisual currently exposes only its Seedance single-first-frame workflow; it has no middle-frame, end-frame or camera-trajectory upload controls. Its four-second, 40-paid-credit option is suitable only when that narrower task fits. Ordinary editing-software keyframes animate properties on a timeline and are outside this guide’s main generation task. No Grok application capability is asserted here.

Sources and next steps

Source check: October 4, 2026. Official contracts, vendor positioning and independent leaderboard observations are labeled separately. No new generation or independent model benchmark was run.

H3 Max reference-to-video contract →

H3 Max image-to-video contract →

H3 Max camera-controls contract →

Utopai X access guide · Keyframe control guide · FLUX 3 Video guide · Seedance limits · PTV single-first-frame alternative · All guides

Checking image studio availability…