Christmas Kitten Snow Scene
A festive video prompt showing a kitten and a woman playing in a gentle snowfall with holiday decorations.
1,188 prompts
A festive video prompt showing a kitten and a woman playing in a gentle snowfall with holiday decorations.
A playful video prompt for dancing in outer space wearing a stylish pink outfit.
A spiritual and cosmic prompt describing a double helix iris formed by galaxies, creating an all-knowing eye.
A complex motion prompt describing the evolution and combination of neon green dragons into a giant frog dragon.
An artistic prompt for a woman in black attire with a decorative flower hat and a mysterious black cat in the background.
A motion generation prompt that takes a static photo of people on a trail and makes them begin riding their mountain bikes.
A highly detailed cinematic prompt for a video showing a man entering a vintage Cadillac in the rain, focusing on atmosphere and tactile movements.
A complex artistic prompt designed to create a tribute to Albert Einstein using reference portraits.
A descriptive prompt for an animation of a character performing among metallic fractals in an inverse realm.
A detailed 10-second cinematic animation prompt where a character catches a high-speed flying coffee mug with athletic precision.
A character animation prompt showing a woman walking forward in a stylish green outfit, designed for smooth movement without obstructions.
A poetic video prompt for a peaceful evening scene involving musical instruments and a soulful, soaring atmosphere.
An extremely detailed multi-shot prompt for a photorealistic vlog-style video of a Korean woman in her bedroom.
A detailed video prompt for a conversation at a cyberpunk noodle bar, including an extension for added interaction.
A transformation video prompt transitioning a puppy into a mystical dire wolf with Fibonacci swirl eyes.
A high-energy animation prompt combining Japanese and Korean advertising styles with a funky theme.
Generates a period-style video of a lady dancing in a beaded dress with silver and pearl accessories
A narrative video prompt for a sci-fi sequence featuring a man discovering a time machine and meeting a robot.
A surreal animation prompt where a dinosaur aids the moon by removing a rocket part through physical action.
A cinematic Game of Thrones prompt featuring Daenerys Targaryen riding her dragon over cliffs at golden hour.
A creative video prompt combining stealth bombers with holiday lights in a pastel color palette.
A video prompt where the camera spins around a pilot taxiing a Boeing 737 toward the runway.
A multi-subject sitcom scene prompt featuring characters from The Office (US) in a meeting room setting.
A surreal video featuring a living, beating golden heart displayed in a luxury foyer surrounded by roses.
Last reviewed August 17, 2026 · Editorial synthesis of current Topview samples and public prompting guidance
Quick answer
A useful Grok Imagine video prompt can be concise: identify the subject, give it one readable action, and say how the camera observes that action. Add the environment, light, visual treatment, dialogue, sound, or constraints only when each detail resolves a production decision. For an image-led workflow, treat the source frame as the visual anchor and describe what should move, what should stay recognizable, and where the action ends. For a more complex clip, arrange two or three short beats in chronological order rather than combining unrelated events in one sentence. Natural language is enough; concrete verbs such as turns, catches, tracks, or pulls back are more directable than labels such as epic or dynamic. If the selected Topview mode provides audio or dialogue generation, name the speaker, exact line, delivery, ambience, and sound cue separately. Finish with a small set of observable realism rules, then review the output for subject drift, motion, framing, physics, speech, and the final frame. Prompt wording can improve clarity, but it cannot guarantee identity, lip sync, sound, or physically perfect motion in every generation.
[output context] + [subject or source-frame anchor] + [primary action or short beats] + [shot and camera] + [environment, light, and style] + [dialogue and sound when available] + [continuity and realism constraints]
Name the intended format, orientation, and clip purpose when they affect composition or pacing. Keep resolution and duration in product settings when the selected workflow exposes those controls.
Define the person, product, animal, vehicle, or environment viewers should follow. In an image-led workflow, say which visible details should remain recognizable without redescribing every pixel.
Use a concrete motion verb and a clear direction, speed, or end state. For a sequence, arrange a small number of actions in the order they should happen.
Choose a shot size, viewpoint, and one motivated camera move. Explain what the move reveals instead of stacking cinematic terms.
Add weather, atmosphere, light direction, palette, and one coherent capture language that supports the action.
When audio is available in the selected Topview mode, separate exact speech from ambience, effects, and music, then connect important sounds to visible events.
Name the few details, screen directions, object states, or physical outcomes that would visibly break the idea if they changed.
Start with the smallest instruction that communicates the idea, then add control only where ambiguity could change the result. Most clips do not need every layer.
| Goal | Prompt pattern | Example wording |
|---|---|---|
| Capture the core idea | [subject] + [one visible action] | A red fox trots across a snow-covered country road. |
| Clarify motion | [subject] + [action direction, pace, and end state] | A red fox trots from frame left to frame right, slows at the tire tracks, and stops to listen. |
| Direct the viewer | [action] + [shot size] + [one camera behavior] | Low side-tracking medium shot follows the fox at its pace, then settles when the fox stops. |
| Set a coherent world | [environment] + [light] + [one visual treatment] + [physical atmosphere] | Quiet rural road at blue hour, cold backlight on the fur, naturalistic documentary texture, loose snow moving in the wind. |
| Protect the result | [important sound if available] + [two or three observable constraints] | For an audio-enabled mode: soft paw steps and winter wind. One fox only, stable leg count, no sudden cut or camera roll. |
The current Topview samples range from one-line actions to detailed shot plans. Use the simplest pattern that still makes the intended event readable.
| Goal | Prompt pattern | Example wording |
|---|---|---|
| Animate a source image | Anchor the existing subject → describe what begins moving → define a restrained end state | Starting from the supplied portrait, she lifts her eyes toward the window, takes one quiet breath, and ends with a slight smile; keep the framing and wardrobe recognizable. |
| Stage a single action | Initial position → one action with direction and weight → clear completion | The cyclist enters from frame right, brakes on the wet pavement with believable momentum, plants one foot, and comes to a complete stop beside the kiosk. |
| Show a transformation | Stable starting form → visible transformation mechanism → finished form and hold | The paper bird unfolds along its creases into a small mechanical swallow; brass joints lock into place, the wings open once, and the finished form holds on the table. |
| Build a short narrative arc | Establish → trigger → reaction or payoff, with one main event per beat | Start wide on the empty platform. A suitcase rolls into view by itself. The waiting conductor notices it, steps back once, and the camera ends on his reaction. |
| Direct a product reveal | Product anchor → controlled material motion → readable hero state | A ribbon of condensation travels down the unchanged bottle while the turntable rotates a quarter turn; the label finishes facing camera in a clean centered hero frame. |
Select camera language according to the information the shot must reveal. One clear move is usually easier to evaluate than several simultaneous moves.
| Goal | Prompt pattern | Example wording |
|---|---|---|
| Reveal emotion | Stable medium shot → slow push-in → stop before an intimate close-up | Medium shot at eye level; slow push-in as the runner hears the announcement, ending on her restrained reaction. |
| Follow movement | Side or rear tracking shot at subject speed + consistent screen direction | Waist-height side-tracking shot follows the skateboarder moving left to right; keep the face near the upper-left third. |
| Show form or scale | Close detail or low angle → controlled orbit or pullback → wider context | Begin on the rover wheel pressing into red dust, then pull back slowly to reveal the vehicle alone beneath the canyon wall. |
| Create grounded realism | Locked or lightly handheld camera + specific imperfection + movement limit | Locked street-level camera with mild autofocus recovery as the bus crosses foreground; no zoom and no camera shake after focus settles. |
| Protect spatial continuity | Declare entrance, travel direction, eyeline, and final screen position | The mug enters from the left, is caught once at center frame, then exits fully to the right; the hand ends empty. |
Use physical details that can be seen or heard. Audio instructions are relevant only when the selected mode supports them, and every output still needs review.
| Goal | Prompt pattern | Example wording |
|---|---|---|
| Make atmosphere visible | [weather or particles] + [how they react to subject and light] | Fine rain streaks through the storefront light, splashes under each step, and beads on the jacket without becoming fog. |
| Keep lighting coherent | [key light source] + [direction] + [surface response] + [continuity rule] | Warm window light remains camera-left, creating one stable rim on the glass and soft reflections across the metal cap. |
| Write short dialogue | [named speaker] says "[exact line]" in [delivery]; [listener or camera behavior] | In an audio-enabled mode, Nia says, "Leave the light on," quietly and without smiling; the listener remains silent in a locked two-shot. |
| Synchronize sound and action | [visible event] lands with [specific effect]; [ambience or music rule] | The station sign flickers out with one electrical pop; distant train ambience continues, with no music. |
| Request believable physics | [weight, contact, momentum, or material behavior] + [observable failure exclusions] | The heavy crate compresses the wet soil on landing, slides only a few centimeters, and stops; no bounce, floating, duplicated crate, or changing dimensions. |
Animate the bicycle courier in a cinematic and realistic way with dramatic camera movement, great sound, and lots of action.
Short 9:16 image-led street scene using the uploaded frame as the visual anchor. Keep the same bicycle courier, yellow rain jacket, black helmet, cargo bag, and red bicycle recognizable. One continuous street-level shot. Start in a locked medium-wide frame as the courier looks over the left shoulder. The bicycle then rolls forward from right to left at a controlled pace while the camera begins a smooth side track. A delivery receipt slips from the cargo bag; the courier brakes once, plants the left foot on the wet pavement, reaches down, and picks it up. End with the bicycle fully stopped and the receipt visible in the courier’s empty right hand. Overcast afternoon, soft storefront reflections, light rain, natural tire spray, believable weight and braking momentum. For an audio-enabled mode: quiet traffic, rain on fabric, one brake squeak, no music or dialogue. Preserve screen direction and bicycle proportions. No cuts, camera roll, extra rider, duplicated wheels, floating receipt, or wardrobe change.
The weak prompt names a mood but does not identify a readable event, camera path, sound cue, or finish state. The stronger version turns the idea into one reviewable action arc, gives the source frame a bounded role, preserves one screen direction, uses a single motivated tracking move, describes contact and momentum, and defines which sounds apply only in an audio-enabled mode. Its constraints target failures that would break this particular shot instead of promising perfect realism or adding a generic negative list.
Add one concrete action, its direction or pace, and a clear end state. In an image-led workflow, focus on what changes after the starting frame.
Choose an observable shot size and move—such as a slow push-in, locked low angle, or side track—and say what it should reveal.
Keep one primary action arc. If the idea needs multiple locations or major actions, divide it into separate generations or a small number of chronological beats.
Use one main camera behavior per beat. Reserve a pan, pullback, or angle change for the moment it adds new information.
State weight, contact, momentum, material response, or environmental interaction, then review the output for the exact failure you care about.
First confirm the selected Topview mode supports the needed audio input or generation. Keep lines short, name the speaker and delivery, and verify synchronization after generation.
List the few visible invariants that matter—face, clothing, product geometry, prop count, light direction, or screen direction—without implying a guarantee.
Describe the intended action positively, then exclude only a few scene-specific failures such as an extra object, an unwanted cut, or a changed product shape.
A good prompt identifies the subject, one readable action, how the camera observes it, and the intended environment or finish. Add sound and constraints only when they matter to the chosen mode and concept. Clear production decisions are more useful than a long string of quality adjectives.
Yes. Many current Topview examples use only a subject and one action. A concise prompt is appropriate when the idea is simple or a source image already establishes the composition. Add detail when you need to control direction, timing, camera, atmosphere, sound, continuity, or the final state.
There is no useful universal word count. Use one or two sentences for a simple movement and a short ordered brief for a more complex action. Remove repeated adjectives, contradictory styles, and instructions that do not change what should move, appear, sound, or remain stable.
Text-to-video needs enough description to establish the subject and scene. Image-to-video can use the starting frame as a visual anchor, so the prompt should concentrate on motion, camera, atmosphere, sound, and the details you want to remain recognizable. Available inputs depend on the Topview mode you select.
Use timestamps or numbered beats when order and pacing would otherwise be ambiguous. For one continuous action, a start, progression, and end state is usually clearer. Multiple cuts increase continuity demands, so split a larger idea into separate clips when each scene needs its own setup.
Begin with familiar instructions such as locked shot, close-up, wide shot, slow push-in, pullback, pan, orbit, side track, handheld follow, POV, or low angle. Choose one main move and connect it to what the audience should notice.
When the selected Topview mode supports audio, write short exact dialogue, identify the speaker and delivery, and separate speech from ambience, effects, and music. Attach important sound cues to visible actions, and review lip sync and timing rather than assuming they will be exact.
Describe observable physics: where weight lands, what touches the ground, how momentum is absorbed, how fabric or particles react, and what state the subject ends in. Keep the action load manageable and evaluate contact, proportions, and object count after generation.
Use a clear source-frame or subject anchor, repeat only the most important visual invariants, limit unnecessary cuts and style changes, and keep product geometry or wardrobe descriptions stable. These steps can improve direction but do not guarantee identical details across every frame.
Start with positive instructions for the intended action, framing, materials, and end state. Then add a few exclusions tied to visible risks in that scene. Whether a separate negative-prompt field is available depends on the selected Topview workflow.
How this guide was built
This guide is an editorial synthesis of a snapshot of 1,188 Grok Imagine records in the Topview prompt library reviewed on August 17, 2026; every record in that snapshot was classified as video generation. About seven in ten non-empty prompts used 40 words or fewer, while longer examples added camera or shot direction, ordered beats, dialogue or sound, preservation language, and scene-specific realism constraints when the concept required more control. The structure was also checked against official xAI material describing image-led motion, camera, pacing, atmosphere, physics, sound design, and current Video 1.5 audio features, plus public Grok Imagine guides covering prompt length, motion verbs, camera vocabulary, use-case examples, mistakes, and FAQs. All formulas and examples above were written from scratch for Topview. The official 1.5 material was used to validate general vocabulary, not to claim that every Topview Grok Imagine mode exposes the same inputs or behavior. These recommendations are starting points, not a guarantee that any prompt has been independently tested or will reproduce the same identity, dialogue, sound, or motion on every run.