Install
$ agentstack add skill-calesthio-generative-media-skills-video-to-audio-foley ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Video-to-audio Foley
Use this skill to turn a silent, generated, animated, game, ad, social, or edit-ready video clip into a credible sound-effects bed. Treat video-to-audio (V2A) as a production accelerator, not as a finished mix. The job is to decide what should be automated, what must be hand-spotted, how to prompt and segment the model, how to preserve sync, and how to verify the result against the picture.
Boundaries
Use automated V2A when:
- the clip is short enough for the available model or can be divided into coherent beats;
- the target is a believable first-pass Foley bed, social clip, animatic, ad temp track, game prototype, or atmospheric pass;
- the visible action has obvious sound causes such as water, footsteps, vehicles, impacts, machinery, crowd, fabric, animals, weather, UI gestures, or prop handling;
- the user can accept iteration and editorial repair.
Do not rely on automated V2A alone when:
- frame-exact sync is mission-critical, such as rhythm games, trailers with hard cuts, slapstick hits, weapon impacts, or product demos where every click matters;
- the output must be delivered as clean reusable game assets or separated stems;
- the source video contains private people, unreleased product footage, client-confidential material, or copyrighted audio you are not allowed to upload to a hosted provider;
- the desired sound is an identifiable real person's voice, a copyrighted sound logo, a famous film/game sound, or a protected musical cue;
- the user needs final dialogue, lip-sync, narration, mix mastering, or music composition rather than Foley/sound effects.
If the user asks for "add sound to this video," first clarify whether they want:
- a single mixed audio track attached to the video;
- separate ambience / Foley / impacts / props / music stems;
- isolated reusable sound assets for a game or library;
- only a prompt/workflow, not generated media.
What is documented versus heuristic
Facts below are source-verified as of 2026-07-10 unless a current project/tool registry says otherwise. Provider endpoints, model IDs, costs, licenses, duration limits, output formats, and commercial-use terms are volatile; re-check live docs before making paid calls or promising a deliverable.
Documented facts:
- Google DeepMind described V2A as research that combines video pixels with natural-language prompts to produce synchronized soundtracks, including sound effects, dialogue, and score-like audio. Do not assume public API access unless current tooling confirms it.
- fal
fal-ai/controlfoleyis documented as a hosted video-to-audio model for synchronized sound effects shaped by text prompts. Its API acceptsvideo_url, optionalprompt,negative_prompt, optional 2-4 secondreference_audio_url,duration, inference controls, and seed; it returns a video with audio and a 44.1 kHz mono WAV audio file. - fal
fal-ai/thinksound/audiois documented as a hosted V2A endpoint that generates realistic audio from a video with an optional text prompt; if no prompt is provided, the prompt may be extracted from the video. Its schema includesvideo_url,prompt,seed,num_inference_steps, andcfg_scale, and returns an audio file plus the prompt used. - MMAudio is a CVPR 2025 open-source model/repo that generates synchronized audio from video and/or text; its README says the default output/training duration is 8 seconds, larger deviations may reduce quality, the default CLI model is
large_44k_v2, and the authors tested on Ubuntu. - FoleyCrafter is an open-source video-to-audio framework that uses a text-to-audio base model with a semantic adapter and temporal controller; it supports text prompt and negative prompt control in its repo examples.
- HunyuanVideo-Foley is a Tencent open-source text-video-to-audio framework; its repo documents single-video and batch inference with
infer.py, an audio description prompt, and XL/XXL model sizes. Check current license and model terms before commercial use. - ElevenLabs Sound Effects is text-to-sound, not video-conditioned V2A. It is useful for manual Foley layers and ambience loops; docs describe duration control from 0.1 to 30 seconds, looping for seamless atmospheres, and prompt-influence control.
- FoleyBench frames V2A Foley quality as both semantic alignment with visible events and temporal alignment with event timing. Use both dimensions in QA.
Production heuristics:
- Prefer one broad V2A pass for continuous ambience and low-risk motion, then layer manual text-to-SFX or library effects for important transients.
- Split clips by acoustic scene, not by arbitrary equal lengths. A cut from subway platform to kitchen should be two generations even if the model accepts the full duration.
- If the model cannot accept explicit timecodes, put the timing discipline in the edit: generate candidates, align transients in the DAW/NLE, trim, crossfade, and layer.
- A good prompt describes the acoustic world, material, perspective, density, and exclusions. It is not just a list of visible objects.
- For social video, "believable and not distracting" usually beats maximum realism. For game assets and hero ads, separation, repeatability, and legal custody matter more than one-click convenience.
Source-video preparation
Before generation:
- Duplicate the source video and preserve the original.
- Confirm the user has rights to upload/process the footage and to publish synthetic audio with it.
- Remove or mute existing audio unless it is intentionally used as a reference and the provider permits it.
- Trim to the natural acoustic unit: one action, one location, one camera beat, or one ambience zone.
- Keep handles of 6-12 frames where possible so crossfades do not cut off attacks or tails.
- Export a clean review file with stable frame rate, visible action, no burn-in UI that changes timing, and the same duration as the intended audio.
- Note frame rate, timecode start, duration, target platform, and desired deliverable format.
- If using a hosted provider that requires public URLs, upload only approved footage; avoid personal data, watermarked client content, unreleased products, or sensitive locations.
For clips longer than the model's reliable duration, segment by scene/action and plan overlap:
- 0.25-0.5 seconds overlap for ambience or crowds;
- 2-6 frames around hard impacts where transients must align;
- no overlap when a cut intentionally changes the acoustic space abruptly.
Spot the video before prompting
Create a spot list even if the model can analyze the video. The spot list is the contract between picture, generation, edit, and QA.
Minimum fields:
- timecode in / out;
- visible cause or implied off-screen cause;
- sound role: ambience, footstep, cloth, prop, impact, vehicle, creature, UI, whoosh, room tone, crowd, mechanical, water, weather, musical sting;
- sync priority: hero, supporting, background, optional;
- acoustic perspective: close mic, camera perspective, distant, muffled, underwater, indoor reflective, outdoor open, phone speaker, helmet cam;
- production method: V2A pass, text-to-SFX, library, manual Foley recording, silence;
- notes on rights, taste, and exclusions.
Example spot list:
| Timecode | Visual event | Sound role | Priority | Method | |---|---|---:|---:|---| | 00:00.000-00:08.000 | Wet alley, neon signs, light rain | rain bed, city hum | supporting | V2A or looped ambience | | 00:01.420 | Boot enters puddle | splash, leather creak | hero | generated SFX + manual sync | | 00:03.100-00:05.700 | Coat swings while walking | cloth movement | background | V2A if natural, otherwise subtle library/recorded cloth | | 00:06.040 | Metal door slams shut | impact, tail reverb | hero | separate SFX; align transient |
Mark silence intentionally. Not every visible movement needs sound; over-Foley makes AI clips feel fake.
Prompt construction
Build prompts from these layers:
- Scene bed: place, room size, weather, crowd, machine tone, distance.
- Hero actions: the 1-4 events the viewer must notice.
- Materials and mechanics: rubber on concrete, glass clink, leather creak, metal scrape, plastic button, wet gravel.
- Perspective and mix density: camera-perspective, close and dry, distant and reverberant, subtle background, no music.
- Exclusions: avoid narration, speech, melody, unrelated animals, wind roar, extra explosions, crowd applause, copyrighted motifs.
- Timing hints only if the provider/model/tool accepts or benefits from them. If not, use short segments and edit sync manually.
Useful wording:
- "camera-perspective production audio"
- "subtle natural Foley, not exaggerated cartoon sounds"
- "dry close-up prop handling with small room reflections"
- "continuous low city ambience under sparse footsteps"
- "single hard ceramic impact at the cut, no music, no voice"
- "avoid extra off-screen events"
Avoid:
- "make it cinematic" without naming the actual acoustic events;
- "perfect sync" as a prompt-only promise;
- overloading a single generation with every micro-movement;
- asking for copyrighted sound-alikes, celebrity voices, or branded sonic logos;
- letting a model add music when the job is Foley.
Provider and model selection
Always inspect the current tool registry, provider docs, licenses, pricing, input schema, and safety terms before choosing. Then pick based on the production need:
| Need | Good fit | Watchouts | |---|---|---| | Hosted V2A with text and negative prompt control | fal ControlFoley or current equivalent | Requires upload/URL; API details and commercial label are volatile; reference audio must be cleared. | | Hosted V2A that can infer a prompt from video | fal ThinkSound or current equivalent | Generated prompt may misread intent; review and override the prompt for brand work. | | Local/open-source reproducibility | MMAudio, FoleyCrafter, HunyuanVideo-Foley, ThinkSound local, or current open model | Check GPU/OS requirements, model license, checkpoint provenance, and duration assumptions. | | High-priority transient or isolated prop | text-to-SFX, recorded Foley, or licensed library effect | Needs manual spotting and sync; may be better than V2A for clean stems. | | Continuous ambience loop | text-to-SFX loop tool, library ambience, or V2A bed | Verify seamless looping and avoid audible repetition/pumping. | | Generated video with native audio | an audiovisual video-generation model | Not the same as adding Foley to an existing edit; less post control and harder to preserve picture lock. |
Decision rule: choose the least magical tool that gives enough sync and control. If a simple text-to-SFX footstep plus a timeline nudge will beat a full V2A pass, use the simple path.
Recommended workflow
- Define deliverables.
- final muxed video only, audio WAV, stems, SFX pack, project file, or all of these;
- target platform loudness/format;
- whether music, dialogue, or voiceover already exists.
- Prepare and segment the video.
- create silent working copies;
- trim by acoustic scenes;
- record exact durations and frame rate;
- decide if hosted upload is allowed.
- Make the spot list.
- identify hero transients first;
- mark ambience separately from Foley;
- note sounds that are implied but not visible;
- mark anything that should remain silent.
- Generate a first V2A bed.
- use one scene-level prompt per segment;
- set seed and parameters if available;
- request no music/no voice unless intentionally desired;
- save raw outputs unchanged.
- Review against picture.
- mute/unmute with the video;
- mark drift, wrong sources, excess events, and missing hero hits;
- decide whether to regenerate, segment further, or replace with manual SFX.
- Layer precision sounds.
- generate or source isolated hits/foley for hero events;
- align transients at the frame or waveform level;
- shape with fades, EQ, reverb, pitch, and volume automation.
- Build the mix.
- keep V2A ambience low enough that manual hero sounds read clearly;
- duck ambience under dialogue/voiceover;
- use crossfades at segment joins;
- avoid clipping and uncontrolled low-frequency buildup.
- QA and package.
- export preview video and separate audio/stems as requested;
- preserve the source video, spot list, prompts, seeds, model IDs, provider, generation date, license notes, and raw generations;
- document any manual edits and rejected generations.
Sync and editing guidance
For hero impacts, sync by transients, not by prompt. A hit should land on the visual contact frame or the frame the edit implies as contact. If the sound blooms late, cut earlier attack from another take or layer a small click/thump on the contact frame.
For footsteps, do not chase every foot if the legs are small or obscured. Sync the first clear step, establish cadence, then let room tone and cloth sell the motion. If the model adds too many steps, use a narrower segment or replace with manual footsteps.
For cloth, keep it quiet and close. Loud cloth movement reads amateur unless the shot is an extreme close-up or animation exaggeration.
For props, name material and action: "ceramic mug placed on wooden desk," "plastic latch clicks twice," "paper envelope crinkles in hand." Generic "object sound" prompts produce mush.
For vehicles and machinery, separate continuous engine/room tone from discrete events such as ignition, brake squeal, gear shift, door slam, pass-by, or shutdown.
For UI/product demos, generated Foley is usually risky. Use designed clicks, taps, whooshes, and interface ticks as separate assets so they match brand taste and exact timing.
For animation, decide the sound world before generating. A realistic foley bed on a stylized cartoon may feel wrong; a hyperreal product sound on a flat motion graphic may feel premium.
Rights, consent, and safety
Before upload or publication:
- verify source-video rights, talent releases, client permissions, and any location/product restrictions;
- verify the provider's current commercial-use and retention terms;
- treat reference audio as copyrighted unless the user proves clearance;
- do not clone, imitate, or synthesize a real person's voice without consent;
- do not recreate famous sound trademarks, sonic logos, game weapon sounds, film creature calls, or identifiable music cues;
- avoid adding deceptive sounds to news, legal, medical, security, or evidence footage unless clearly labeled as reconstruction or dramatization;
- preserve provenance so downstream editors know what is synthetic, generated, licensed, recorded, or client-provided.
If the user asks to add sounds that change the factual meaning of real footage, state the risk and propose labeling, abstaining, or making a clearly fictional edit.
QA checklist
Semantic fit:
- Are the audible sources plausible for what is visible or intentionally implied?
- Did the model hallucinate animals, voices, crowds, music, machinery, or weather?
- Are material cues believable: metal, glass, cloth, water, wood, rubber, skin, plastic?
Temporal fit:
- Do hero hits land on contact frames?
- Do footsteps match cadence well enough for shot size?
- Is there drift across the segment?
- Do cuts and transitions reset the acoustic space cleanly?
Acoustic fit:
- Does reverb match the location and camera distance?
- Is the bed too dry, too wet, too loud, too clean, or too noisy?
- Do close sounds feel attached to the camera perspective?
- Are layers fighting in the same frequency range?
Technical fit:
- No clipping, crackle, dropouts, double attacks, codec smear, or obvious loops.
- Exports are the requested sample rate, channels, codec, and duration.
- Final video and audio durations match; no tail is cut unless intentional.
- Stems and project assets are named and stored with provenance.
Editorial fit:
- The audio supports the story rather than showing off generation.
- Important visual events are
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: calesthio
- Source: calesthio/generative-media-skills
- License: MIT
- Homepage: https://github.com/calesthio/OpenMontage
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.