QueyFramesQueyFrames
gemini aikeyframesaudioblogearly access
‹ blog

How to write AI video prompts that control pacing

Text-to-video models will give you the shot and invent the timing. Here is how to write the timing into the prompt — naming beats in order, pinning them to seconds, and letting wired keyframes place themselves in the sentence.

13 August 20265 min read

contents

  1. Why the timing falls out of the prompt
  2. Three habits that do most of the work
  3. Let the keyframes write that part for you
  4. A worked example
  5. What is not worth writing
  6. Read the prompt before you pay for it

Most people writing their first text-to-video prompt describe a subject, a style and a mood, get back a clip that looks roughly right, and then find they have no idea how to make the next one land a beat earlier. The description was never the hard part. Pacing is the part prompts lose, and it is the part you notice when the clip has to sit next to a cut, a lyric or a sound effect.

Why the timing falls out of the prompt

Look at what a current video model actually accepts as a request. With Gemini's video generation, the only creative knob outside the prompt text is the aspect ratio. Duration, frame rate and resolution are fixed — a clip is three to ten seconds at 720p and 24fps, and the model decides where in that range it lands. There is no negative prompt, no temperature, no system instruction and no per-second control track.

So everything you want to say about when has to live in the words. That is not a limitation to work around; it is the actual interface. A prompt that reads like a shot list gets shot-list pacing back, and a prompt that reads like a mood board gets whatever pacing the model felt like.

Three habits that do most of the work

  • Name the beats in order. Not "a drone shot of a coastline" but "opens tight on the waterline, pulls back to reveal the headland, settles on the lighthouse". Three clauses, three beats, in the order they happen.
  • Pin them to seconds. "At 0:02 the camera starts pulling back" gives the model a place to put the change. Vague ordering gets vague ordering back; a number gets something close to the number.
  • End on a state, not an action. The last second of a generated clip is where things go strange, because the model has nothing to resolve toward. "...and holds on the lighthouse" is a cheap way to buy yourself a usable final frame — which matters if the next clip continues from it.

None of that is specific to one product. It is just what happens when the only timing channel you have is prose, so you use prose deliberately.

Let the keyframes write that part for you

Writing seconds by hand works until the timing comes from somewhere else — the end of a chorus, a cut you already made, a section you marked on a waveform. Then you are retyping numbers that already exist somewhere in the project, and they go stale the moment the source moves.

In QueyFrames that is what the keyframe ports are for. Wire a keyframe track into a video node and every keyframe is numbered across the input slots, then placed into the prompt individually by token: {keyframe1}, {keyframe2}, and so on. You write the sentence; the timings drop into it.

  1. Get the timings from wherever they really live — mark them on a waveform, take the section ends the music node publishes, or drop them on a clip's own ruler.
  2. Wire that track into the video node's keyframe slots. A slot takes any number of edges, so a keyframes node fanned out one segment per port still lands as one ordered list.
  3. Write the prompt around the tokens: "opens on the empty stage, {keyframe1}, then {keyframe2} and holds".
  4. Anything you don't place explicitly merges into one block at {keyframes}, or is appended at the end if you never name it.

Each keyframe writes itself as its description if it has one, or as the properties it moves ([0:00–0:03] opacity 0 → 1) if it doesn't, or as a bare cut point if it is only a mark. A fade is still something a model can act on; collapsing it to "cut here" throws away the instruction.

note

Keyframes past the clip's maximum length are dropped, not clamped. Crushing three late beats onto the last frame produces a worse clip than not mentioning them, and it produces one that is hard to debug — you would be reading a prompt that says three things happen at 0:10.

A worked example

The version that gets you a slot machine

"Cinematic shot of a lighthouse on a rocky coast at dusk, dramatic clouds, moody, 35mm film look."

Everything here is style. There is not one word about sequence, so every re-roll gives you a different edit of the same material and none of them cut against anything.

The version you can cut with

"Dusk, 35mm film look. Opens tight on waves breaking over rock. At 0:02 the camera begins a slow pull back, revealing the headland. At 0:05 the lighthouse beam sweeps once across frame. Settles and holds on the lighthouse against the clouds."

Same style words, but now there are four beats with places to be. Re-roll this and the clips differ in texture rather than in structure, which is what makes the second attempt useful instead of just different.

What is not worth writing

  • Negative prompts. "No text, no watermark, no distortion" has no request field behind it here, and spending prompt on what you don't want costs you room to say what you do.
  • Resolution and frame rate. Fixed by the model. Asking for 4K 60fps in prose gets you nothing except a longer prompt.
  • Exact durations. You can influence where in the three-to-ten-second range a clip lands by how much you ask it to do, but you cannot request 6.5 seconds. Cut to length afterwards on a timeline instead.

Read the prompt before you pay for it

Once a prompt is assembled from wired inputs, tokens and your own text, the thing that gets sent is no longer the thing you typed. Generation costs credits, and a surprise is an expensive way to discover you left a token unplaced — so hover a wired field and the resolved prompt is in its title, exactly as the request will carry it. Checking takes a second and does not cost a generation.

If you want the model underneath this, how QueyFrames uses Gemini for video covers text-to-video and image-to-video and what each one is actually good for. If you want the timing side, keyframes as shared timing is the piece that makes {keyframe1} possible in the first place.

Keep reading

  • 6 August 20264 min read

    Image to video with AI: first frames, reference images, and the difference

    An image handed to a video model can be two completely different instructions — the frame it starts on, or the look it borrows. Here is how to tell them apart, why the order you send them in matters, and how to wire a still into a clip.

  • 30 July 20264 min read

    How to sync AI music to your video, in both directions

    Cutting picture to a track and writing a track to a cut are the same problem read from opposite ends. Both come down to one thing: a timing that means the same instant to every tool that reads it.

  • 24 August 20265 min read

    How to keyframe AI animation: spans, not points

    Point keyframes only mean something beside their neighbors. When an AI reads your timing, a span — start, end, what changes — stands alone. How to write them.

Build the graph these posts are about.

QueyFrames is in early access. Create an account to join the list.

sign up for early access→

QueyFrames — AI video generator by QuokkaQuery

overviewgemini aikeyframesaudioblogtermsprivacysign up