Install
$ agentstack add skill-aiagentwithdhruv-skills-image-to-video ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Image-to-Video AI Generation — Skill Reference
> Version: 1.0.0 | Updated: 2026-03-02 | Category: Content & Video
Table of Contents
- [Tool Comparison Matrix](#1-tool-comparison-matrix)
- [Detailed Tool Profiles](#2-detailed-tool-profiles)
- [Universal Prompt Best Practices](#3-universal-prompt-best-practices)
- [Camera Movement Reference](#4-camera-movement-reference)
- [Subject Animation Guide](#5-subject-animation-guide)
- [Consistency & Stability](#6-consistency--stability)
- [Avoiding Distortion & Common Mistakes](#7-avoiding-distortion--common-mistakes)
- [Text & Logo Preservation](#8-text--logo-preservation)
- [Thumbnail-to-Video Specific Guide](#9-thumbnail-to-video-specific-guide)
- [Prompt Templates](#10-prompt-templates)
1. Tool Comparison Matrix
| Tool | Best Model (Mar 2026) | Max Length | I2V | Free Tier | Max Resolution | Native Audio | Best For | |------|----------------------|------------|-----|-----------|----------------|-------------|----------| | Runway | Gen-4.5 | 10s | Yes | 125 one-time credits (~25s Gen-4 Turbo) | 4K (upscale) | No | Cinematic consistency, character ref | | Kling | Kling 3.0 / 2.6 Pro | 15s (3.0) / 10s (2.6) | Yes | 66 daily credits (360-540p, watermark) | 1080p (Master) | Yes (2.6+) | Motion control, product detail, fashion | | Pika | Pika 2.5 | 10s | Yes | 80 monthly credits (480p, watermark) | 1080p+ (paid) | No | Creative effects (Pikaswaps, Pikadditions) | | Luma | Ray3 / Ray3 Modify | 20s (720p+) | Yes | 30 gens/month (draft res, watermark) | 1080p | No | Long clips, start+end frame, cinematic | | Sora | Sora 2 / Sora 2 Pro | 25s | Yes | None (Plus $20/mo minimum) | 1080p (Pro: 1792x1024) | Yes | Narrative scenes, physics, dialogue | | Vidu | Vidu Q3 | 16s | Yes | 3 videos/month (720p, watermark) | 4K (Q3 Pro) | Yes (native) | Multi-shot sequences, synced audio | | Hailuo/MiniMax | Hailuo 2.3 | 10s | Yes | Daily bonus credits (720p, watermark) | 1080p (paid) | Yes (2.6+) | Speed, social content, A/B testing | | Google Veo | Veo 3.1 | 8s | Yes (Ingredients) | Limited (Gemini free: older model) | 4K (3840x2160) | Yes | 4K output, film language, camera control | | Adobe Firefly | Firefly Video | 5s | Yes | Limited credits (with CC sub) | 2K native (up to 8K upscale) | No | Commercial-safe (IP indemnity), integration with CC | | Seedance | Seedance 2.0 (ByteDance) | 15s | Yes | Free credits on signup | 1080p | Yes (native) | Multimodal input, fast generation | | WAN | WAN 2.6 / 2.1 | 10s | Yes | Open source (run locally) | 1080p | No | Open source, self-hosted, general-purpose |
2. Detailed Tool Profiles
Runway (Gen-4 / Gen-4.5)
Current Models:
- Gen-4.5 (latest, Jan 2026): State-of-the-art motion quality, prompt adherence, visual fidelity. Variable durations 2-10s.
- Gen-4 Turbo: Fast, economical (5 credits/sec vs Gen-4.5 at 25 credits/sec). Good for iteration.
- Gen-4: Mid-tier (12 credits/sec). Balanced quality/cost.
Image-to-Video Specifics:
- Upload a reference image + text prompt describing motion
- Choose duration (5 or 10 seconds) and aspect ratio
- Enable "Fixed Seed" for reproducible motion
- Reference images maintain character appearance, clothing, features across scenes
- Strong spatial understanding — objects/backgrounds stay coherent during camera movement
Pricing:
- Free: 125 one-time credits (~25s of Gen-4 Turbo, ~5s of Gen-4.5)
- Standard: $12-15/mo (625 credits)
- Pro: $28-35/mo (2,250 credits)
- Unlimited: $76-95/mo (2,250 fast + unlimited relaxed)
Prompt Best Practices (Runway-Specific):
- Focus prompts EXCLUSIVELY on motion — do NOT re-describe what is in the image
- Start simple, iterate by adding detail
- Use camera terms: pan, tilt, dolly, orbit, zoom, truck, pedestal, crane, rack focus, crash zoom
- Structure: "The camera [motion] as [subject action]"
- Abstract/conceptual language causes unpredictable results — be specific and physical
- Re-describing image elements in detail can reduce motion or cause artifacts
Sources: Runway Pricing | Gen-4 Research | Gen-4.5 Research
Kling AI (Kling 3.0 / 2.6 Pro)
Current Models:
- Kling 3.0 (latest): Scene-aware generation, character/prop consistency, native audio, 3-15s clips
- Kling 2.6 Pro: Built-in English and Chinese audio, stronger prompt control, cinematic realism
- Kling 2.6 Motion Control: Upload a motion reference video to guide character movement
- Variants: Turbo (fast), Pro (balanced), Master (highest quality)
Image-to-Video Specifics:
- Upload image as subject + describe movement in prompt
- Motion Control mode: image + reference video for precise motion transfer
- Preserves edges, logos, and fabric details (great for product/fashion)
Pricing:
- Free: 66 daily credits (resets every 24h, no rollover). 360-540p, watermarked, non-commercial
- Paid plans: $6.99-180/mo depending on tier
Prompt Best Practices (Kling-Specific):
- For I2V: describe ONLY what should move/change + camera behavior. The image IS the scene.
- Keep ONE main action ("hero action"). Hint at secondary motion only.
- For Motion Control: do NOT describe motion in prompt (the reference video defines it). Use prompt for environment/look only.
- Use terms like "slow push-in", "drone follow", "lateral track"
- Describe pace with words like "glides smoothly" or "jerks to a halt"
- Ensure character limbs are visible in source image (hidden limbs cause hallucination/extra fingers)
- Leave "breathing room" around subject for movement
- Match aspect ratios between image and motion reference
Sources: Kling AI | Kling 3.0 Guide | Kling 2.6 Motion Control
Pika Labs (Pika 2.5)
Current Models:
- Pika 2.5 (latest): Sharper, smoother cinematic clips. Upgraded engine.
- Pikaformance: Talking face model for lifelike voice-to-face performances
- AI Selves: Personalized AI avatar creation
Key Features:
- Pikaframes: Turn 2-5 images into smooth transition video with realistic movement
- Pikaswaps: Replace objects in video (e.g., dog -> robot) with preserved lighting/motion
- Pikadditions: Insert new characters/objects into footage
- Scene Ingredients: Upload your own characters/objects for consistency
Pricing:
- Free: 80 monthly credits. 480p only, watermarked, non-commercial
- Paid: Unlocks all resolutions, removes watermark, commercial use
Prompt Best Practices (Pika-Specific):
- Great for creative/stylized transformations rather than photorealistic
- Use Pikaframes for multi-image storytelling
- Specify lighting and physics behavior for realistic material interactions
- Best for short creative social clips and effects-heavy content
Sources: Pika Pricing | Pika 2.5 Release
Luma Dream Machine (Ray3)
Current Models:
- Ray3: Primary generation model. Supports 5-20s video depending on resolution.
- Ray3 Modify: Modify existing footage with character reference images
- Ray3.14: Draft resolution model (available on free tier)
Image-to-Video Specifics:
- Upload still image, animate with natural motion and cinematic camera action
- Start+End frame feature: provide first and last frame, AI generates the transition
- Adds subtle camera pans, zooms, perspective shifts automatically
Pricing:
- Free: $0/mo. 30 gens/month. Draft resolution (Ray3.14), 720p images, watermarked, personal only
- Lite: $9.99/mo. 3,200 credits. 1080p images, watermarked, non-commercial
- Plus: $29.99/mo. 10,000 credits. No watermark, commercial rights
- Unlimited: $94.99/mo. 10,000 fast + unlimited relaxed
Video Duration by Resolution:
- 540p SDR: 5s (160 credits), 10s (320 credits)
- 720p SDR: 5-20s
- 1080p SDR: up to 20s
Sources: Luma Pricing | Dream Machine | Ray3 Info
OpenAI Sora (Sora 2)
Current Models:
- Sora 2: Text-to-video and image-to-video with synchronized audio
- Sora 2 Pro: Higher resolution (1792x1024) and better quality
Image-to-Video Specifics:
- Start with a still image and expand it into motion
- Physically accurate, realistic, controllable
- Can insert people into any Sora-generated environment with accurate appearance and voice
- Native dialogue and sound effects generation
Pricing:
- NO free tier (as of Jan 10, 2026)
- ChatGPT Plus ($20/mo): Unlimited 480p video generation
- ChatGPT Pro ($200/mo): Higher quality, priority access
- API: $0.10/sec (720p), $0.30/sec (720p Pro), $0.50/sec (1024p Pro)
Prompt Best Practices (Sora-Specific):
- Rewards prompts describing INTENT and MOOD, not just motion
- Use director-style framing and gradual motion introduction
- Structure prompts in distinct sections: what happens, visual style, audio elements
- Be explicit about sound (dialogue, foley, music, mood)
- Specify character positioning, framing, emotional states, gestures
- Describe physics: "gentle collision" vs "violent crash", "heavy object slides" vs "light feather floats"
- Support 15-25 second clips. Describe pacing progression.
- Specify 24fps for cinematic feel
Sources: Sora 2 Guide | Sora Announcement
Vidu (Vidu Q3)
Current Models:
- Vidu Q3 (latest): Native audio+video in one pass, up to 16s, 2K resolution, multi-shot "Smart Cuts"
- Vidu Q2: Previous gen. Natural motion, film-like camera effects.
- Reference-to-Video 2.0: Character/subject consistency across generations
Key Features:
- First AI model to generate multi-shot, edited-style sequences with synced audio from a single prompt
- "Smart Cuts" for automatic multi-shot sequences
- Audio: BGM + SFX synced to scene rhythm
- Up to 4K in Q3 Pro via API
Pricing:
- Free: 3 videos/month. 720p, watermarked
- Paid plans available on vidu.com
Sources: Vidu | Vidu Q3 Guide | Vidu Q3 on WaveSpeed
Hailuo / MiniMax (Hailuo 2.3)
Current Models:
- Hailuo 2.3 (latest): Improved physical actions, stylization, character micro-expressions, anime support
- Hailuo 02: Standard and Fast variants. 768p and 1080p, up to 10s
- Media Agent: Multi-modal creation with minimal manual editing
Pricing:
- Free: $0/mo. Daily bonus credits. 720p, watermarked. Peak-hour wait times.
- Standard: $9.99/mo. 1,000 credits, fast-track, no watermark, up to 5 tasks
- Unlimited: $94.99/mo. Unlimited credits
Prompt Best Practices (Hailuo-Specific):
- Works best with clean images and modest motion requests
- Great for rapid A/B testing and short-form social content
- Strong anime/stylized content support in 2.3
Sources: Hailuo AI | MiniMax Hailuo 2.3
Google Veo (Veo 3.1)
Current Models:
- Veo 3.1: 4K output (3840x2160), vertical video (9:16), "Ingredients to Video" (up to 4 reference images)
- Veo 3 Standard: Older model available to some free users
- Veo 3 Fast: Lower-cost option
Key Features:
- FIRST mainstream AI model with true 4K output
- "Ingredients to Video": Accept up to 4 reference images per generation
- Character identity consistency across scene changes
- Native vertical video for YouTube Shorts / TikTok / Reels
- Built-in audio generation
Pricing:
- Free (Gemini): 100 monthly AI credits for Flow/Whisk. May get Veo 3 Standard (not 3.1)
- Pro ($19.99/mo): Limited Veo 3.1 access
- Ultra ($124.99/3mo or ~$42/mo): 25,000 monthly credits, full Veo 3.1
- API: Veo 2 at $0.35-0.50/sec
Prompt Best Practices (Veo-Specific):
- Excels with film language — reference shot types and pacing
- Separate subject stability from camera motion in prompts
- Input images should be 720p+ with 16:9 or 9:16 aspect ratio
- Prompts referencing specific shot types produce more controlled results
Sources: Veo 3.1 4K Update | Veo 3.1 Blog | Google DeepMind Veo
Adobe Firefly Video
Current Model: Firefly Video (Feb 2026)
Key Features:
- 5s clips per generation
- Native 2K resolution (up to 8K with Upscale)
- IP indemnity — commercially safe, trained on licensed content
- QuickCut: Upload b-roll or generate footage, auto-create structured first cut
- Deep integration with Premiere Pro, After Effects, Creative Cloud
Pricing:
- Firefly Standard: $9.99/mo (2,000 premium credits). ~20 videos at 100 credits/5s clip
- Firefly Pro: $19.99/mo (4,000 premium credits)
- Firefly Premium: $199.99/mo (50,000 premium credits)
- Jan-Mar 2026 promo: Unlimited generations on paid plans
Best For: Enterprise/agency use where IP indemnity matters. Integration with existing Adobe workflows.
Sources: Adobe Firefly Pricing | Firefly Blog
Seedance 2.0 (ByteDance)
Current Model: Seedance 2.0
Key Features:
- Unified multimodal audio-video joint generation (text, image, audio, video inputs)
- 4-15s video length
- 1080p resolution
- 30% faster than Seedance 1.0
- Native audio generation (BGM + SFX)
Pricing:
- Free credits on signup (check-in daily for more)
Sources: Seedance 2.0 | Seedance on fal.ai
WAN 2.6 / 2.1 (Open Source)
Current Models:
- WAN 2.6: Latest release
- WAN 2.1: Widely available, open-source on Hugging Face
Key Features:
- Open source — run locally, no credits needed
- 1.3B and 14B parameter variants
- Text AND image generation in video (Chinese + English)
- Realistic physics simulation
- Great general-purpose all-rounder
Pricing: Free (open source). Hardware costs only.
Best For: Self-hosted workflows, privacy-sensitive projects, unlimited generation without credits
Sources: WAN GitHub | WAN on HuggingFace
3. Universal Prompt Best Practices
The Golden Rules
- Separate identity from motion. The image defines WHO/WHAT. The prompt defines HOW it MOVES.
- Do NOT re-describe the image. This causes reduced motion or visual artifacts.
- Start simple, iterate. Begin with one action, one camera move. Add complexity after testing.
- Be physically specific, not conceptual. "Camera slowly pushes in" > "dramatic emphasis"
- 3-4 descriptive elements per component is the sweet spot. More adjectives past this degrades quality.
Prompt Structure Formula
[Camera movement], [pace/speed], [subject action], [environmental motion/details]
Example:
Slow push-in, steady cinematic pace, the developer's fingers type on the glowing keyboard,
holographic UI panels float and pulse with soft blue light around the workspace
The 8-Point Shot Grammar (Advanced)
For consistent cinematic outputs, cover these 8 elements:
| Element | What to Specify | Example | |---------|----------------|---------| | 1. S
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: aiagentwithdhruv
- Source: aiagentwithdhruv/skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.