Grok Imagine Video 1.5
xAI's image-to-video model. Takes a static image and brings it to life with realistic motion and built-in audio.
Last updated
xAI's image-to-video model. Takes a static image and brings it to life with realistic motion and built-in audio.
Fast and affordable, it suits both quick iteration and high-volume production for simple to advanced use cases.
Director-level camera control: Understands cinematic language — pan, tilt, dolly, orbit, tracking, aerial, handheld, push-in.
Content-aware motion: Adapts to the input — exaggerated physics for illustrated characters, 360° rotation for products, natural expressions for portraits.
Native synchronized audio: Visual and sound produced together — music, effects, ambience, and brief dialogue — with no separate audio editing step.
Speed and cost: Affordable enough to iterate freely and to run at volume.
Product showcases: Turn a product shot into a 360° rotation or a hero turn with dramatic lighting.
Video ad first frames: Animate a static opening frame into a compelling motion intro.
Character & mascot animation: Bring illustrated brand characters to life with smooth motion.
Social content: Short clips with sound for Reels, TikTok, and feeds, in native vertical or square ratios.
Image-to-video only — every run needs an input image. For text-to-video, use Grok Imagine Video (non-1.5).
720p ceiling — no 1080p or 4K.
Stability falls off with length — 5–8s is reliable; 15s clips are more prone to artifacts.
Preview release — behavior and availability may change.
The model already sees your image — prompt for motion, not description.
Key principles
Describe the action, camera move, and atmosphere — not what's already in the frame.
Always specify a shot type and camera movement.
Use specific verbs with intensity modifiers ("racing past at high speed," not "car passing").
Negative prompts are ignored — describe what you want instead.
Keep it to one subject, one action, one camera move; iterate in small steps.
Last updated

