The reusable formula
Start with: subject and starting state + primary motion + environment motion + camera direction + visual treatment + timing. Treat it as a checklist rather than a sentence you must fill with adjectives.
- Subject: who or what the shot follows.
- Primary motion: one clear action with direction and pace.
- Environment: the motion happening around the subject.
- Camera: framing, movement and final composition.
- Treatment: lighting, texture and genre only where they change the result.
- Timing: how the action develops across the available clip.
Text-to-video example
A cyclist waits beneath a rain-soaked flyover at night. She pushes off and accelerates toward camera while traffic reflections stretch across the road. The camera tracks backwards at handlebar height, then eases into a wide side profile. Natural wheel spray, sodium-vapour light, grounded documentary realism.
Image-to-video example
The camera slowly pushes toward the subject as she turns from the window and meets the lens. Curtains move gently in the background. Keep the expression restrained and preserve the room's existing light.
For image-to-video, the reference already describes appearance and composition. Follow Runway's current guidance and spend the prompt on motion, camera and temporal progression instead of redundantly describing every visible object.
Why overloaded prompts fail
- Several competing actions force the model to choose what to ignore.
- Contradictory camera instructions create unstable framing.
- Long style lists dilute the visual priority.
- Demanding multiple scenes from a short single-shot generation compresses every beat.
Iterate one variable at a time
Save the first output, identify the single largest miss, and change only the clause responsible for it. If motion is wrong, do not simultaneously change lighting, lens and subject. afterSora projects make this easier by keeping prompts, references and outputs together for comparison.
Final pre-generation check
- Is there one unmistakable primary action?
- Does the camera instruction have one start and one finish?
- Does the requested action fit the chosen duration?
- For image-to-video, are you describing motion instead of restating the image?
- Have you selected the right aspect ratio before generating?
