AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Image To Video

skill-aiagentwithdhruv-skills-image-to-video · by aiagentwithdhruv

Generate AI video from static images using Kling 3.0, Hailuo, Luma Ray3, Runway Gen-4.5, and 8 other tools. Covers free vs paid tools, prompt writing (motion-only), camera control, and face stability. Use when user asks to animate an image, create AI video, or convert photo to video.

No reviews yet
0 installs
45 views
0.0% view→install

Install

$ agentstack add skill-aiagentwithdhruv-skills-image-to-video

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-aiagentwithdhruv-skills-image-to-video)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Image To Video? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Image-to-Video AI Generation — Skill Reference

> Version: 1.0.0 | Updated: 2026-03-02 | Category: Content & Video


Table of Contents

  1. [Tool Comparison Matrix](#1-tool-comparison-matrix)
  2. [Detailed Tool Profiles](#2-detailed-tool-profiles)
  3. [Universal Prompt Best Practices](#3-universal-prompt-best-practices)
  4. [Camera Movement Reference](#4-camera-movement-reference)
  5. [Subject Animation Guide](#5-subject-animation-guide)
  6. [Consistency & Stability](#6-consistency--stability)
  7. [Avoiding Distortion & Common Mistakes](#7-avoiding-distortion--common-mistakes)
  8. [Text & Logo Preservation](#8-text--logo-preservation)
  9. [Thumbnail-to-Video Specific Guide](#9-thumbnail-to-video-specific-guide)
  10. [Prompt Templates](#10-prompt-templates)

1. Tool Comparison Matrix

| Tool | Best Model (Mar 2026) | Max Length | I2V | Free Tier | Max Resolution | Native Audio | Best For | |------|----------------------|------------|-----|-----------|----------------|-------------|----------| | Runway | Gen-4.5 | 10s | Yes | 125 one-time credits (~25s Gen-4 Turbo) | 4K (upscale) | No | Cinematic consistency, character ref | | Kling | Kling 3.0 / 2.6 Pro | 15s (3.0) / 10s (2.6) | Yes | 66 daily credits (360-540p, watermark) | 1080p (Master) | Yes (2.6+) | Motion control, product detail, fashion | | Pika | Pika 2.5 | 10s | Yes | 80 monthly credits (480p, watermark) | 1080p+ (paid) | No | Creative effects (Pikaswaps, Pikadditions) | | Luma | Ray3 / Ray3 Modify | 20s (720p+) | Yes | 30 gens/month (draft res, watermark) | 1080p | No | Long clips, start+end frame, cinematic | | Sora | Sora 2 / Sora 2 Pro | 25s | Yes | None (Plus $20/mo minimum) | 1080p (Pro: 1792x1024) | Yes | Narrative scenes, physics, dialogue | | Vidu | Vidu Q3 | 16s | Yes | 3 videos/month (720p, watermark) | 4K (Q3 Pro) | Yes (native) | Multi-shot sequences, synced audio | | Hailuo/MiniMax | Hailuo 2.3 | 10s | Yes | Daily bonus credits (720p, watermark) | 1080p (paid) | Yes (2.6+) | Speed, social content, A/B testing | | Google Veo | Veo 3.1 | 8s | Yes (Ingredients) | Limited (Gemini free: older model) | 4K (3840x2160) | Yes | 4K output, film language, camera control | | Adobe Firefly | Firefly Video | 5s | Yes | Limited credits (with CC sub) | 2K native (up to 8K upscale) | No | Commercial-safe (IP indemnity), integration with CC | | Seedance | Seedance 2.0 (ByteDance) | 15s | Yes | Free credits on signup | 1080p | Yes (native) | Multimodal input, fast generation | | WAN | WAN 2.6 / 2.1 | 10s | Yes | Open source (run locally) | 1080p | No | Open source, self-hosted, general-purpose |


2. Detailed Tool Profiles

Runway (Gen-4 / Gen-4.5)

Current Models:

  • Gen-4.5 (latest, Jan 2026): State-of-the-art motion quality, prompt adherence, visual fidelity. Variable durations 2-10s.
  • Gen-4 Turbo: Fast, economical (5 credits/sec vs Gen-4.5 at 25 credits/sec). Good for iteration.
  • Gen-4: Mid-tier (12 credits/sec). Balanced quality/cost.

Image-to-Video Specifics:

  • Upload a reference image + text prompt describing motion
  • Choose duration (5 or 10 seconds) and aspect ratio
  • Enable "Fixed Seed" for reproducible motion
  • Reference images maintain character appearance, clothing, features across scenes
  • Strong spatial understanding — objects/backgrounds stay coherent during camera movement

Pricing:

  • Free: 125 one-time credits (~25s of Gen-4 Turbo, ~5s of Gen-4.5)
  • Standard: $12-15/mo (625 credits)
  • Pro: $28-35/mo (2,250 credits)
  • Unlimited: $76-95/mo (2,250 fast + unlimited relaxed)

Prompt Best Practices (Runway-Specific):

  • Focus prompts EXCLUSIVELY on motion — do NOT re-describe what is in the image
  • Start simple, iterate by adding detail
  • Use camera terms: pan, tilt, dolly, orbit, zoom, truck, pedestal, crane, rack focus, crash zoom
  • Structure: "The camera [motion] as [subject action]"
  • Abstract/conceptual language causes unpredictable results — be specific and physical
  • Re-describing image elements in detail can reduce motion or cause artifacts

Sources: Runway Pricing | Gen-4 Research | Gen-4.5 Research


Kling AI (Kling 3.0 / 2.6 Pro)

Current Models:

  • Kling 3.0 (latest): Scene-aware generation, character/prop consistency, native audio, 3-15s clips
  • Kling 2.6 Pro: Built-in English and Chinese audio, stronger prompt control, cinematic realism
  • Kling 2.6 Motion Control: Upload a motion reference video to guide character movement
  • Variants: Turbo (fast), Pro (balanced), Master (highest quality)

Image-to-Video Specifics:

  • Upload image as subject + describe movement in prompt
  • Motion Control mode: image + reference video for precise motion transfer
  • Preserves edges, logos, and fabric details (great for product/fashion)

Pricing:

  • Free: 66 daily credits (resets every 24h, no rollover). 360-540p, watermarked, non-commercial
  • Paid plans: $6.99-180/mo depending on tier

Prompt Best Practices (Kling-Specific):

  • For I2V: describe ONLY what should move/change + camera behavior. The image IS the scene.
  • Keep ONE main action ("hero action"). Hint at secondary motion only.
  • For Motion Control: do NOT describe motion in prompt (the reference video defines it). Use prompt for environment/look only.
  • Use terms like "slow push-in", "drone follow", "lateral track"
  • Describe pace with words like "glides smoothly" or "jerks to a halt"
  • Ensure character limbs are visible in source image (hidden limbs cause hallucination/extra fingers)
  • Leave "breathing room" around subject for movement
  • Match aspect ratios between image and motion reference

Sources: Kling AI | Kling 3.0 Guide | Kling 2.6 Motion Control


Pika Labs (Pika 2.5)

Current Models:

  • Pika 2.5 (latest): Sharper, smoother cinematic clips. Upgraded engine.
  • Pikaformance: Talking face model for lifelike voice-to-face performances
  • AI Selves: Personalized AI avatar creation

Key Features:

  • Pikaframes: Turn 2-5 images into smooth transition video with realistic movement
  • Pikaswaps: Replace objects in video (e.g., dog -> robot) with preserved lighting/motion
  • Pikadditions: Insert new characters/objects into footage
  • Scene Ingredients: Upload your own characters/objects for consistency

Pricing:

  • Free: 80 monthly credits. 480p only, watermarked, non-commercial
  • Paid: Unlocks all resolutions, removes watermark, commercial use

Prompt Best Practices (Pika-Specific):

  • Great for creative/stylized transformations rather than photorealistic
  • Use Pikaframes for multi-image storytelling
  • Specify lighting and physics behavior for realistic material interactions
  • Best for short creative social clips and effects-heavy content

Sources: Pika Pricing | Pika 2.5 Release


Luma Dream Machine (Ray3)

Current Models:

  • Ray3: Primary generation model. Supports 5-20s video depending on resolution.
  • Ray3 Modify: Modify existing footage with character reference images
  • Ray3.14: Draft resolution model (available on free tier)

Image-to-Video Specifics:

  • Upload still image, animate with natural motion and cinematic camera action
  • Start+End frame feature: provide first and last frame, AI generates the transition
  • Adds subtle camera pans, zooms, perspective shifts automatically

Pricing:

  • Free: $0/mo. 30 gens/month. Draft resolution (Ray3.14), 720p images, watermarked, personal only
  • Lite: $9.99/mo. 3,200 credits. 1080p images, watermarked, non-commercial
  • Plus: $29.99/mo. 10,000 credits. No watermark, commercial rights
  • Unlimited: $94.99/mo. 10,000 fast + unlimited relaxed

Video Duration by Resolution:

  • 540p SDR: 5s (160 credits), 10s (320 credits)
  • 720p SDR: 5-20s
  • 1080p SDR: up to 20s

Sources: Luma Pricing | Dream Machine | Ray3 Info


OpenAI Sora (Sora 2)

Current Models:

  • Sora 2: Text-to-video and image-to-video with synchronized audio
  • Sora 2 Pro: Higher resolution (1792x1024) and better quality

Image-to-Video Specifics:

  • Start with a still image and expand it into motion
  • Physically accurate, realistic, controllable
  • Can insert people into any Sora-generated environment with accurate appearance and voice
  • Native dialogue and sound effects generation

Pricing:

  • NO free tier (as of Jan 10, 2026)
  • ChatGPT Plus ($20/mo): Unlimited 480p video generation
  • ChatGPT Pro ($200/mo): Higher quality, priority access
  • API: $0.10/sec (720p), $0.30/sec (720p Pro), $0.50/sec (1024p Pro)

Prompt Best Practices (Sora-Specific):

  • Rewards prompts describing INTENT and MOOD, not just motion
  • Use director-style framing and gradual motion introduction
  • Structure prompts in distinct sections: what happens, visual style, audio elements
  • Be explicit about sound (dialogue, foley, music, mood)
  • Specify character positioning, framing, emotional states, gestures
  • Describe physics: "gentle collision" vs "violent crash", "heavy object slides" vs "light feather floats"
  • Support 15-25 second clips. Describe pacing progression.
  • Specify 24fps for cinematic feel

Sources: Sora 2 Guide | Sora Announcement


Vidu (Vidu Q3)

Current Models:

  • Vidu Q3 (latest): Native audio+video in one pass, up to 16s, 2K resolution, multi-shot "Smart Cuts"
  • Vidu Q2: Previous gen. Natural motion, film-like camera effects.
  • Reference-to-Video 2.0: Character/subject consistency across generations

Key Features:

  • First AI model to generate multi-shot, edited-style sequences with synced audio from a single prompt
  • "Smart Cuts" for automatic multi-shot sequences
  • Audio: BGM + SFX synced to scene rhythm
  • Up to 4K in Q3 Pro via API

Pricing:

  • Free: 3 videos/month. 720p, watermarked
  • Paid plans available on vidu.com

Sources: Vidu | Vidu Q3 Guide | Vidu Q3 on WaveSpeed


Hailuo / MiniMax (Hailuo 2.3)

Current Models:

  • Hailuo 2.3 (latest): Improved physical actions, stylization, character micro-expressions, anime support
  • Hailuo 02: Standard and Fast variants. 768p and 1080p, up to 10s
  • Media Agent: Multi-modal creation with minimal manual editing

Pricing:

  • Free: $0/mo. Daily bonus credits. 720p, watermarked. Peak-hour wait times.
  • Standard: $9.99/mo. 1,000 credits, fast-track, no watermark, up to 5 tasks
  • Unlimited: $94.99/mo. Unlimited credits

Prompt Best Practices (Hailuo-Specific):

  • Works best with clean images and modest motion requests
  • Great for rapid A/B testing and short-form social content
  • Strong anime/stylized content support in 2.3

Sources: Hailuo AI | MiniMax Hailuo 2.3


Google Veo (Veo 3.1)

Current Models:

  • Veo 3.1: 4K output (3840x2160), vertical video (9:16), "Ingredients to Video" (up to 4 reference images)
  • Veo 3 Standard: Older model available to some free users
  • Veo 3 Fast: Lower-cost option

Key Features:

  • FIRST mainstream AI model with true 4K output
  • "Ingredients to Video": Accept up to 4 reference images per generation
  • Character identity consistency across scene changes
  • Native vertical video for YouTube Shorts / TikTok / Reels
  • Built-in audio generation

Pricing:

  • Free (Gemini): 100 monthly AI credits for Flow/Whisk. May get Veo 3 Standard (not 3.1)
  • Pro ($19.99/mo): Limited Veo 3.1 access
  • Ultra ($124.99/3mo or ~$42/mo): 25,000 monthly credits, full Veo 3.1
  • API: Veo 2 at $0.35-0.50/sec

Prompt Best Practices (Veo-Specific):

  • Excels with film language — reference shot types and pacing
  • Separate subject stability from camera motion in prompts
  • Input images should be 720p+ with 16:9 or 9:16 aspect ratio
  • Prompts referencing specific shot types produce more controlled results

Sources: Veo 3.1 4K Update | Veo 3.1 Blog | Google DeepMind Veo


Adobe Firefly Video

Current Model: Firefly Video (Feb 2026)

Key Features:

  • 5s clips per generation
  • Native 2K resolution (up to 8K with Upscale)
  • IP indemnity — commercially safe, trained on licensed content
  • QuickCut: Upload b-roll or generate footage, auto-create structured first cut
  • Deep integration with Premiere Pro, After Effects, Creative Cloud

Pricing:

  • Firefly Standard: $9.99/mo (2,000 premium credits). ~20 videos at 100 credits/5s clip
  • Firefly Pro: $19.99/mo (4,000 premium credits)
  • Firefly Premium: $199.99/mo (50,000 premium credits)
  • Jan-Mar 2026 promo: Unlimited generations on paid plans

Best For: Enterprise/agency use where IP indemnity matters. Integration with existing Adobe workflows.

Sources: Adobe Firefly Pricing | Firefly Blog


Seedance 2.0 (ByteDance)

Current Model: Seedance 2.0

Key Features:

  • Unified multimodal audio-video joint generation (text, image, audio, video inputs)
  • 4-15s video length
  • 1080p resolution
  • 30% faster than Seedance 1.0
  • Native audio generation (BGM + SFX)

Pricing:

  • Free credits on signup (check-in daily for more)

Sources: Seedance 2.0 | Seedance on fal.ai


WAN 2.6 / 2.1 (Open Source)

Current Models:

  • WAN 2.6: Latest release
  • WAN 2.1: Widely available, open-source on Hugging Face

Key Features:

  • Open source — run locally, no credits needed
  • 1.3B and 14B parameter variants
  • Text AND image generation in video (Chinese + English)
  • Realistic physics simulation
  • Great general-purpose all-rounder

Pricing: Free (open source). Hardware costs only.

Best For: Self-hosted workflows, privacy-sensitive projects, unlimited generation without credits

Sources: WAN GitHub | WAN on HuggingFace


3. Universal Prompt Best Practices

The Golden Rules

  1. Separate identity from motion. The image defines WHO/WHAT. The prompt defines HOW it MOVES.
  2. Do NOT re-describe the image. This causes reduced motion or visual artifacts.
  3. Start simple, iterate. Begin with one action, one camera move. Add complexity after testing.
  4. Be physically specific, not conceptual. "Camera slowly pushes in" > "dramatic emphasis"
  5. 3-4 descriptive elements per component is the sweet spot. More adjectives past this degrades quality.

Prompt Structure Formula

[Camera movement], [pace/speed], [subject action], [environmental motion/details]

Example:

Slow push-in, steady cinematic pace, the developer's fingers type on the glowing keyboard,
holographic UI panels float and pulse with soft blue light around the workspace

The 8-Point Shot Grammar (Advanced)

For consistent cinematic outputs, cover these 8 elements:

| Element | What to Specify | Example | |---------|----------------|---------| | 1. S

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.