Topview
    • MCP/Skill
    • Plugin
    • API
    • Pricing
    1. Home
    2. Grok Imagine Video 1 5

    AI Ads

    AI Video AgentAI Ads VideoAI Product VideoAI UGC VideoURL to Video

    AI Avatar

    AI Avatar GeneratorProduct AvatarDesign My AvatarAI Lip Sync

    AI Video

    AI Video GeneratorDrama StudioAI Video Body SwapAI Video UpscalerAI Video Watermark RemoverTikTok Watermark RemoverSora Watermark RemoverSubtitle RemoverAI Motion ControlAI Lip SyncURL to Video

    AI Image

    AI Image GeneratorAI Face SwapAI Image Character SwapAI Image UpscalerAI InpaintAI Image Text EditorPhoto Angle EditorAI RelightingAI Product PhotographyAI Virtual Try-OnAI Storyboard GeneratorExtract Color PaletteImage to Prompt

    AI Voice

    VoiceoverAI Voice CloningAI Music Generator

    3D

    3D World Generator3D Shot Composer

    Use Cases

    AdvertisingAffiliate MarketingEcommerceDTC BrandsAI Live Stream

    Resources

    BlogAffiliate ProgramLearning CenterAlternativeAPI

    Company

    Topview StudioPrivacy PolicyTerms
    Topview

    © 2026 TOPVIEW PTE. LTD.

    Singapore: 20 Collyer Quay #20-03, Singapore 049319
    Los Angeles: 15970 Los Serranos CC Dr #251, Chino Hills, CA 91709

    Grok Imagine Video 1.5 — xAI's #1Image-to-Video Model with Native Audio

    Generate cinematic 720p/24fps videos up to 15 seconds from text prompts and reference images. Native audio — dialogue, sound effects, and ambience — generated in a single pass with the Aurora autoregressive engine.

    Grok Imagine Video 1.5
    Grok Imagine Video 1.5
    Barefoot businessman skateboarding in San Francisco — reference image
    @Image1
    428/3500

    Animate the scene starting from @Image1@Image1. The barefoot blonde businessman in the navy suit skateboards rapidly down a steep San Francisco street in a dynamic low-angle tracking shot. He speaks enthusiastically with arms wide, looking directly into the camera — hair and tie fluttering in the wind. He then jumps and grinds a metal handrail alongside the sidewalk, landing cleanly. Cinematic camera tracking, realistic motion, 8K.

    Coming Soon

    See What Grok Imagine Video 1.5 Can Create

    From native audio-driven cinematic dialogue to anime motion, commercial transitions, and fantasy worldbuilding — explore the kind of stunning videos Grok Imagine Video 1.5 can generate from text prompts and reference images with 720p/24fps output.

    Native Audio & Speech — All in One Pass

    Dialogue, sound effects, and ambience are generated together with video — not dubbed in later. Speech lands on the action, clearer and better synced.

    Prompt

    1950s hotel elevator. A woman in an emerald gown speaks to the operator in a red uniform as the gold-trimmed doors close. Soft dramatic lighting, rich film colors.

    Try Now

    Dynamic Commercial-to-Set Transitions

    Create engaging behind-the-scenes and transition-focused marketing content. Grok Imagine Video 1.5 smoothly transitions from pristine commercial product shots to complex studio sets with fully synchronized foley.

    Prompt

    A continuous pull-out shot. A hand pours milk into a mason jar of iced coffee next to a stack of cookies. The camera pulls back dynamically, revealing a woman taking a bite of the cookie with a synchronized crunch, on a busy green-screen production studio set with crew.

    Try Now

    Stylized Anime & Motion Consistency

    Render vibrant anime art and fluid character motion. Grok Imagine Video 1.5 maintains flawless character details, complex fabric physics, and expressive facial acting across dynamic stylized shots.

    Prompt

    Stylized anime 3D animation. A cute blue-haired elf girl with red eyes, wearing a black cyberpunk Qipao with a blue dragon print and tactical straps, dances playfully in front of a traditional temple with red lanterns. Smooth fluid movement, expressive winks and smiles, vibrant lighting, highly detailed.

    Try Now

    Multi-Agent Physics & Interactions

    Simulate hyper-realistic animal locomotion and chaotic city physics. Grok Imagine Video 1.5 seamlessly coordinates natural animal movement, flocking bird dynamics, and volumetric steam in crowded environments.

    Prompt

    A rabbit sprinting through NYC, fast-paced, photorealistic.

    Try Now

    Narrative Character Growth & Worldbuilding

    Deliver continuous character evolution and epic world-scale transitions. Grok Imagine Video 1.5 simulates biological growth — like hatching and aging — while maintaining character identity across vast, physics-rich fantasy environments.

    Prompt

    A cinematic fantasy sequence. A cute white baby dragon hatches from a shimmering, iridescent egg surrounded by glowing crystals. The dragon grows and spreads its wings on a cliffside, then takes off to fly smoothly through fluffy clouds. It transitions into soaring majestically toward a breathtaking sunset over a vast landscape of floating islands and giant crystal spires.

    Try Now

    Grok Imagine Video 1.5 vs Seedance 2.0: AI Video Model Comparison

    Both Grok Imagine Video 1.5 and Seedance 2.0 are top-tier image-to-video models with native audio, but they serve different priorities. Grok Imagine Video 1.5 prioritizes generation speed and single-pass audio-visual coherence. Seedance 2.0 prioritizes reference depth and multi-shot control.

    Grok Imagine Video 1.5

    Grok Imagine Video 1.5 — built on Aurora autoregressive (MoE) engine. Generates 720p video with native audio in a single pass at ~25s for a 6s clip.

    Seedance 2.0

    Seedance 2.0 — Dual Branch Diffusion Transformer. Excels at multi-shot storytelling with broader reference input support and 1080p output.

    Comparison PointGrok Imagine Video 1.5Seedance 2.0Key Difference
    DeveloperxAIByteDanceDifferent research teams and architectures
    ArchitectureAurora autoregressive (MoE)Dual Branch Diffusion TransformerGrok uses autoregressive; Seedance uses diffusion-based generation
    Generation Speed (6s clip)~25s (Fast mode)~120sGrok Imagine Video 1.5 is ~5× faster
    Max Resolution720p1080pSeedance offers higher max resolution
    Max Duration15s15sBoth support up to 15-second clips
    Native Audio OutputSingle-pass: dialogue, SFX, ambienceDialogue, SFX, lip-syncBoth deliver complete audio-visual generation
    Input TypeImage + Text promptImage + Text + Multi-ref supportSeedance accepts more reference images per generation
    Arena Leaderboard (I2V)#1 (May 2026)#2Grok Imagine Video 1.5 currently leads
    Try on Topview

    Grok Imagine Video — Model Evolution Timeline

    From the launch of xAI's first image-to-video model to the Arena-topping 1.5 — here's how Grok Imagine Video has evolved.

    2025

    Grok Imagine Video 1.0 Launch

    xAI launched its first dedicated image-to-video model — separate from the Grok chatbot. Built on the proprietary Aurora autoregressive engine, it generated up to 10-second 720p clips at 24fps from text and image inputs, quickly gaining traction among creators.

    February 2026

    Multi-Image Support & Extension

    xAI added multi-image support and video extension capabilities to Grok Imagine Video 1.0, allowing creators to chain reference images and extend generated clips for more complex storytelling workflows.

    March 2026

    API Preview & Developer Access

    Grok Imagine Video 1.0 became available via the xAI developer platform API, opening the model to third-party integrations and creative tools like Topview for broader production use.

    Late May 2026

    Grok Imagine Video 1.5 Preview

    xAI released Grok Imagine Video 1.5 in preview. It immediately claimed the #1 position on the Image-to-Video Arena leaderboard with a 52 Elo point jump over version 1.0, surpassing Seedance 2.0 and other competitors. Key upgrades: faster generation (~25s Fast mode), native audio improvements, and extended 15-second clip duration.

    June 16, 2026

    Grok Imagine Video 1.5 Generally Available

    Grok Imagine Video 1.5 exits preview and becomes generally available (GA) on the xAI API. Alongside the GA launch, xAI rolled out new creative workflow features — Projects for organizing work, parallel agent execution for running multiple prompts at once, and search for finding past generations quickly.

    Ongoing

    Ecosystem & Platform Growth

    Grok Imagine Video models are available through the xAI API (grok-imagine-video-1.5), grok.com/imagine web app, iOS and Android apps, and third-party creative platforms. Built on Aurora autoregressive architecture trained on 110,000 NVIDIA GB200 GPUs.

    Now Available

    Model Parameters

    Core Grok Imagine Video 1.5 specifications relevant to creators evaluating output quality, generation speed, and production fit.

    Model

    Grok Imagine Video 1.5

    xAI's latest image-to-video model, launched June 2026

    Arena Ranking

    #1 Image-to-Video

    1404 Elo, +52 over v1.0 (May 2026)

    Audio Output

    Native: dialogue, SFX, ambience, music

    Generated in single pass with video — no post-dubbing

    Engine

    Aurora Autoregressive (MoE)

    Proprietary mixture-of-experts architecture by xAI

    Training Compute

    110,000 GB200 GPUs

    One of the largest GPU clusters for video AI

    Resolution

    480p (draft) / 720p (output)

    24fps cinema-standard frame rate

    Duration

    6s - 15s per clip

    Extendable via chaining

    Aspect Ratios

    7 (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3)

    Full platform coverage from square to vertical

    Input Types

    Image + Text

    Upload a reference image with natural language prompt

    Output Format

    H.264 MP4

    Input accepts JPG, PNG, WEBP, GIF, AVIF

    Generation Speed

    ~25s (6s 720p Fast)

    ~2× faster than v1.0; ~4-5× faster than competitors

    API Endpoint

    grok-imagine-video-1.5

    Generally Available via xAI API

    What's New in Grok Imagine Video 1.5 — v1.0 vs v1.5 Comparison

    Grok Imagine Video 1.5 is xAI's latest image-to-video model, built on the upgraded Aurora autoregressive engine. It delivers faster speeds, better motion physics, improved audio sync, and longer clip durations compared to version 1.0.

    CapabilityGrok Imagine Video 1.0Grok Imagine Video 1.5Improvement
    Max Resolution720p720pBetter detail, less warping
    Generation Speed (6s 720p)~40+ seconds~25 seconds (Fast)Nearly 2x faster
    Max Duration10s15s50% longer clips
    Native AudioBasic syncClearer dialogue, better lip-sync, event-aligned SFXMore polished audio-visual coherence
    Motion PhysicsSome warpingBetter momentum, fewer warps, believable weightMore realistic movement
    Aspect Ratios5 formats7 formats (1:1 to 16:9 and vertical)Full platform-native support
    API StatusPreviewGenerally Available (GA)Production-ready API
    Arena LeaderboardStrong contender#1 Image-to-Video (+52 Elo jump)Top-ranked model
    Platform FeaturesBasic generationProjects, parallel agents, search libraryNew creative workflow tools
    Training ComputePrevious cluster110,000 GB200 GPUsMassive infrastructure scale

    Grok Imagine Video 1.5 vs Seedance 2 vs Veo 4 vs Sora 2 - Model Comparison

    Choosing the right AI video model in 2026 means comparing output quality, speed, and workflow fit. This comparison focuses on the features that matter most for creators, marketers, and production teams.

    FeatureGrok Imagine Video 1.5Seedance 2Veo 4Sora 2
    DeveloperxAIByteDanceGoogleOpenAI
    Max Duration15s15s20s+12s
    Max Resolution720p1080p4K1080p
    Native AudioDialogue + SFX + ambience (single-pass)Dialogue + SFX + lip-syncDialogue + ambience mixGenerated audio
    Input TypeImage + TextImage + Text + Multi-refImage + TextImage + Text
    ArchitectureAurora autoregressive (MoE)Dual Branch Diffusion TransformerDiffusion TransformerDiffusion Transformer
    Generation Speed~25s (6s 720p Fast)~2 min~2.5 min~3 min
    Multi-Shot SequencesVia chainingYesYesYes
    Arena Ranking (I2V)#1 (May 2026)#2Top 5Top 5
    API AvailableGA (grok-imagine-video-1.5)FullFullLimited
    Best ForFast I2V with native audio, rapid iterationReference depth and multi-shot storytellingCinematic polish and 4K outputPhysics realism and text-to-video

    Grok Imagine Video 1.5 stands out as the fastest image-to-video model with native audio — generating high-quality 720p clips in about 25 seconds, roughly 4-5× faster than competitors. It ranked #1 on the Image-to-Video Arena leaderboard as of May 2026. For creators prioritizing speed, native audio-visual coherence, and production efficiency, Grok Imagine Video 1.5 is the clear frontrunner.

    Explore more AI video models on Topview

    Who Should Use Grok Imagine Video 1.5 on Topview

    Grok Imagine Video 1.5 is built for teams that need fast image-to-video generation with native audio — from cinematic storytellers and product marketing teams to social content creators.

    Filmmakers and Story-First Creators

    When you need cinematic framing, camera language, and scene composition from a reference image, Grok Imagine Video 1.5's Aurora engine delivers coherent motion and native audio in about 25 seconds — fast enough for creative exploration.

    Fashion, Beauty, and Product Teams

    Start from a product photo and generate polished product showcase videos. Grok Imagine Video 1.5 excels at maintaining product detail and lighting mood from the reference image with realistic motion and ambiance.

    Performance Marketers and Ad Teams

    Grok Imagine Video 1.5's ~25-second Fast mode makes it ideal for ad variant testing. Generate multiple hooks, angles, and versions rapidly — compare performance and scale what works without slowing down your creative pipeline.

    Music and Dance Creators

    Native audio-visual sync means beat-aware motion and rhythm-driven visuals. Generate performance clips that match music energy without external audio alignment work — all in a single generation pass.

    Viral Social and Trend Creators

    Grok Imagine Video 1.5's speed makes it perfect for social-first creators who need trending hooks, pet videos, and POV concepts at the pace of platform culture. 720p is the sweet spot for social platforms.

    Creative Teams That Value Speed

    If your bottleneck is generation speed, Grok Imagine Video 1.5's 25-second Fast mode is a significant advantage. More iterations, more variants, more chances to find the creative that performs.

    How to Use Grok Imagine Video 1.5

    Upload image and prompt for Grok Imagine Video 1.5
    Step 1

    Upload a reference image and write a prompt

    Start with your key visual — a product photo, character design, or scene reference. Describe the motion, camera movement, and audio atmosphere you want.

    Generate video with Grok Imagine Video 1.5
    Step 2

    Generate Video

    Click generate and watch Grok Imagine Video 1.5 create a 720p/24fps video with native audio in about 25 seconds (Fast mode).

    Download video from Grok Imagine Video 1.5
    Step 3

    Download the video

    Export a clean MP4 with synchronized audio when you're ready to publish to any platform.

    Start Creating Now

    Experience Grok Imagine Video 1.5 — The #1 Image-to-Video AI

    No expensive GPUs required. Generate cinema-grade 720p video with native audio from text prompts and reference images — all in about 25 seconds with Grok Imagine Video 1.5 on Topview.

    Generate with Grok Imagine Video 1.5 Now

    Start free · No credit card required · All leading AI video models in one workspace

    Frequently Asked Questions

    Grok Imagine Video 1.5 is xAI's latest image-to-video AI model, released as Generally Available in June 2026. Built on the Aurora autoregressive (Mixture-of-Experts) engine, it generates 720p/24fps video up to 15 seconds from text prompts and reference images, with native synchronized audio — dialogue, sound effects, and ambience — all in a single inference pass.

    Grok Imagine Video 1.5 uses xAI's proprietary Aurora autoregressive engine — a fundamentally different architecture from the diffusion-based approaches used by Sora and Veo. This enables tightly coupled audio-visual generation in a single pass, rather than generating video first and dubbing audio later. It also delivers significantly faster generation: ~25 seconds for a 6-second 720p clip (Fast mode), roughly 4-5× faster than comparable models.

    Grok Imagine Video 1.5 is significantly faster (~25s vs ~120s for a 6s clip) and currently ranked #1 on the Image-to-Video Arena leaderboard. Seedance 2.0 offers higher max resolution (1080p vs 720p) and broader reference input support. Both models generate native audio alongside video. Choose Grok for speed and audio-visual coherence; Seedance for maximum resolution and reference depth.

    Yes — this is one of its standout features. The Aurora engine generates dialogue, sound effects, ambient audio, and music in the same inference pass as the video. This single-pass approach eliminates external audio stitching and ensures better timing alignment between sound and visuals.

    Grok Imagine Video 1.5 generates clips from 6 to 15 seconds. Multiple clips can be chained together for longer sequences, and the API supports granular duration control at any integer second from 1 to 15.

    You provide a reference image (JPG, PNG, WEBP, GIF, or AVIF) and a natural language prompt describing the motion, camera movement, and audio atmosphere you want. The model generates the video and audio together from these inputs. Grok Imagine Video 1.5 is primarily an image-to-video model — it does not accept video or audio files as input.

    Grok Imagine Video 1.5 is available through xAI's API with usage-based pricing. On Topview, you can access Grok Imagine Video 1.5 alongside other leading AI video models. Topview offers a free tier to get started without a credit card. Check xAI's official API documentation and Topview's pricing page for current rates.

    Grok Imagine Video 1.5 Fast generates a 6-second 720p clip in approximately 25 seconds — nearly 2× faster than the previous model and roughly 4-5× faster than competing models like Seedance 2.0. A 10-second clip typically takes 35-40 seconds.

    No — Grok Imagine Video 1.5 is an image-to-video model. It accepts a reference image and a text prompt, then generates the output video and audio. It does not accept video clips or audio files as input.

    Grok Imagine Video 1.5 supports 7 aspect ratios: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, and 2:3 — covering everything from square Instagram posts to widescreen YouTube to vertical TikTok and Shorts formats.

    Topview offers a free tier that lets you try Grok Imagine Video 1.5 without a credit card. This makes it easy to test the model, compare outputs against other AI video models, and validate your creative direction before committing to a paid plan.

    Grok Imagine Video 1.5 is significantly faster (~25s vs ~3 min) and includes native audio generation in a single pass. Sora 2 excels at physics realism and pure text-to-video quality but lacks native audio and operates at a slower generation speed. Grok Imagine Video 1.5 is better suited for rapid iteration; Sora 2 for high-fidelity physics simulation.

    Veo 4 offers higher max resolution (4K vs 720p) and strong camera control, positioning it as a premium cinematic option. Grok Imagine Video 1.5 competes on speed (~25s vs ~2.5 min) and native single-pass audio generation — making it the practical choice for teams that prioritize iteration speed and audio-visual coherence over peak resolution.

    Grok Imagine Video 1.5 currently outputs at up to 720p resolution, which is well-suited for social media, short-form content, and fast creative iteration — the primary use cases it was designed for. For 4K production needs, models like Veo 4 may be more appropriate.

    Commercial use depends on your Topview plan and xAI's API terms of service. Always confirm usage rights before deploying outputs in ads, client projects, product pages, or other commercial campaigns. Topview's paid tiers are designed for commercial production workflows.

    xAI is the AI company founded by Elon Musk that develops Grok Imagine Video. Based in the United States, xAI builds large-scale AI models including the Grok chatbot and Grok Imagine Video. Their video models run on the Aurora autoregressive engine, trained on a massive 110,000 NVIDIA GB200 GPU cluster — one of the largest AI training infrastructures in the world.

    Third-party model and brand names are trademarks of their respective owners. Topview provides access to supported AI models through its independent platform and is not the developer of those third-party models.

    Grok Imagine Video 1.5 is developed by xAI and was released as Generally Available on June 16, 2026. Specifications are sourced from xAI's official announcements (x.ai/news), the Image-to-Video Arena leaderboard, and verified technical documentation. Arena ranking as of May 2026.

    Best For
    Fast I2V, native audio, rapid iteration
    Reference depth, multi-shot, 1080p output
    Grok for speed; Seedance for reference variety