# Localization Dubbing Production

> Provider-independent localization and dubbing production direction for AI agents producing translated videos, dubbed ads, social clips, explainers, avatar videos, documentaries, training media, and multi-market campaigns. Use when planning or reviewing subtitle-versus-dub strategy, translation adaptation, glossary/style-guide setup, casting or voice matching, lip-sync or phrase-sync direction, sy…

- **Type:** Skill
- **Install:** `agentstack add skill-calesthio-generative-media-skills-localization-dubbing-production`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [calesthio](https://agentstack.voostack.com/s/calesthio)
- **Installs:** 0
- **Category:** [Developer Tools](https://agentstack.voostack.com/c/developer-tools)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [calesthio](https://github.com/calesthio)
- **Source:** https://github.com/calesthio/generative-media-skills/tree/main/skills/production/audio-craft/localization-dubbing-production
- **Website:** https://github.com/calesthio/OpenMontage

## Install

```sh
agentstack add skill-calesthio-generative-media-skills-localization-dubbing-production
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Localization and dubbing production

Use this skill to turn a source video into market-ready localized variants without treating translation as a word swap. Direct the localization strategy, script adaptation, voice approach, timing, accessibility, rights checks, review loop, and quality control before choosing any provider or model.

This is production guidance, not legal advice. When the work involves talent contracts, union-covered performers, biometric/voice data, likeness rights, political content, regulated claims, minors, health/finance/legal claims, or paid media in multiple jurisdictions, require qualified legal/local market review before release.

## Evidence posture

Documented facts in this skill are grounded in the sources listed at the end and were verified on 2026-07-10 unless otherwise stated. Empirical observations are recurring production patterns seen in localization/dubbing work and should be checked against the actual source media. Production heuristics are decision rules for agents; override them when a client brief, platform specification, law, accessibility requirement, or native reviewer says otherwise.

## Start by classifying the job

Identify the content type before making localization choices:

- Entertainment, documentary, testimonial, news, or creator content: preserve voice, identity, cultural specificity, and trust. Subtitles may be more honest than dubbing for testimony or journalism.
- Advertisement, product launch, app demo, social clip, avatar video, or explainer: optimize comprehension, retention, brand tone, and platform behavior. Dubbing or native re-voicing is often worth the cost.
- Training, compliance, medical, finance, legal, or public-sector media: prioritize accuracy, terminology control, accessibility, traceable review, and auditable approvals over fluency shortcuts.
- Children/family, literacy-limited, or sound-off mobile audiences: do not assume subtitles alone are sufficient; consider dubbing, captions, visual reinforcement, and simpler reading load.
- High-emotion drama, close-up presenter, avatar, or face-heavy ad: decide early whether lip-sync is required, because it changes the script, casting, edit, and QC burden.

Ask for or infer these essentials:

1. Source language, target locales, audience, platform, release date, and region.
2. Existing contracts/permissions for voice, likeness, music, on-screen people, and synthetic processing.
3. Source assets: final video, clean dialogue/music/effects stems if available, transcript, captions, brand guide, approved product terms, campaign claims, legal disclaimers, and edit project or shot list.
4. Delivery mode per locale: subtitles, SDH/captions, voice-over, phrase-sync dub, lip-sync dub, localized avatar performance, or fully re-edited market cut.
5. Reviewers: native-language reviewer, brand/legal reviewer, subject-matter reviewer, accessibility reviewer, and final approver.

## Separate facts, observations, and heuristics during planning

Documented facts to carry into production:

- WCAG 2.1 requires captions for prerecorded audio in synchronized media at Success Criterion 1.2.2 Level A and live captions at 1.2.4 Level AA. Audio description is separately required at 1.2.5 Level AA when visual information is not otherwise available.
- The FCC describes TV closed-caption quality around accuracy, synchronicity, completeness, and placement. Even when FCC rules do not apply to a given web/social deliverable, these dimensions are useful caption QC categories.
- Netflix timed text general requirements specify subtitle event duration minimum 5/6 second and maximum 7 seconds for Netflix deliveries; do not apply Netflix-specific limits blindly to non-Netflix platforms, but use them as a reality check for readability.
- ISO 17100:2015 defines requirements for core translation processes and resources for translation services; its public abstract states raw machine-translation output plus post-editing is outside its scope.
- YouTube multi-language audio lets eligible creators upload additional language audio tracks to one video or Short; uploaded tracks must be audio-only and roughly the same length as the video. YouTube automatic dubbing is separate and may contain errors in pronunciation, dialect, idioms, proper nouns, jargon, speech recognition, or voice matching.
- YouTube requires creators to disclose realistic AI-generated or meaningfully AI-altered content in Studio; YouTube says disclosure is not required for caption creation and lists cloning one's own voice for voice overs or dubs among examples creators do not need to disclose, but this is a platform policy, not a blanket rights clearance.
- WebVTT is a W3C text-track format for time-aligned captions/subtitles, descriptions, chapters, and metadata.
- NAVA defines active consent for synthetic/AI voices as an affirmative waiver/agreement by the performer or representative, excluding all other use. SAG-AFTRA public AI resources emphasize consent, disclosure, and digital replica protections in union contexts.

Empirical observations to verify on each project:

- Literal translations often break timing, humor, on-screen UI correspondence, mouth movement, legal nuance, and brand tone.
- AI dubbing systems are most likely to fail on overlapping speakers, noisy audio, regional accents, names, acronyms, jokes, sarcasm, code-switching, rapid speech, songs, and domain jargon.
- Viewers tolerate looser sync in narration, training, explainers, and off-camera voice-over than in close-up human performance, testimonials, drama, or avatar spokespeople.
- Some locales expect dubbing as a default for entertainment; others prefer subtitles for authenticity. Treat this as market research, not a universal rule.

Production heuristics:

- Localize intent first, wording second, timing third. A fluent line that misses the claim, joke, or call to action is wrong.
- Lock the source edit before localization unless the localization plan explicitly includes per-market recuts.
- Build a termbase before translating more than a sample. Fixing terminology after recording is expensive.
- Use native reviewers for market fitness, not just bilingual reviewers for dictionary accuracy.
- Prefer phrase-sync over lip-sync when the mouth is not central to trust or immersion; spend lip-sync budget only where the face sells the performance.
- Treat every synthetic voice, cloned voice, face reenactment, and avatar-lip-sync request as a consent-and-disclosure gate before production, not as a post-delivery paperwork issue.

## Choose subtitles, captions, voice-over, phrase-sync, or lip-sync

Use the lightest mode that meets comprehension, trust, accessibility, and platform goals.

Subtitles are usually strongest when:

- The original performance, testimony, accent, music, or documentary authenticity matters.
- Budget or timeline cannot support native casting, direction, mix, and QC.
- The content is watched with sound on and reading load is reasonable.
- Legal or technical constraints make replacing the original voice risky.
- The source has many speakers but visible lip-sync is not required.

Captions/SDH are required when accessibility is in scope:

- Include dialogue, speaker identification where needed, and meaningful non-speech audio such as music cues, sound effects, offscreen voices, laughter, alarms, or tone-setting audio.
- Captions are not merely translated subtitles; they represent audio information for viewers who cannot hear it.
- For localized accessibility, create captions for the localized audio, not only translated subtitles from the source.

Voice-over is usually strongest when:

- The original speakers may remain audible under a translated narration bed.
- The content is documentary, news, training, webinar, interview, or educational, and face-perfect sync would feel artificial or unnecessary.
- Speed and clarity matter more than the illusion that the person is speaking the target language.

Phrase-sync dubbing is usually strongest when:

- The localized voice should begin and end near the source speaker's phrases, but exact mouth-shape matching is not needed.
- The source contains presenter shots, demos, explainers, corporate training, or ads with moderate face visibility.
- The script can be adapted to match phrase duration while preserving a natural target-language performance.

Lip-sync dubbing is usually strongest when:

- The audience sees close-up mouths, emotional performance, characters, avatars, or direct-to-camera talent.
- The localized version must create the illusion that the on-screen person is speaking the target language.
- The budget supports adaptation, casting, voice direction, retakes, alignment, and visual QC.

Avoid lip-sync when:

- The source is mostly off-camera narration, screen capture, product demo, or motion graphics.
- The target language expansion would force rushed, unnatural speech.
- Consent for face/voice manipulation is missing.
- The production cannot afford native linguistic and performance QC.

## Build the localization brief

Produce or request a brief with:

- Source and target locales, not just languages: e.g. Spanish for Mexico, Spanish for Spain, French for Canada, Arabic MSA versus dialect, Portuguese Brazil versus Portugal.
- Audience: age, expertise, literacy, cultural context, accessibility needs, and likely viewing mode.
- Purpose: inform, persuade, train, entertain, convert, comply, support, or build trust.
- Brand voice: formal/informal, humor tolerance, pronoun/register policy, forbidden tones, approved slogans.
- Market constraints: legal claims, medical/financial disclaimers, regulated terminology, political sensitivity, religious/cultural restrictions, units/currency/date formats.
- Platform/delivery: YouTube multi-language audio, social burned-in captions, OTT timed-text package, LMS training module, broadcast, paid ad upload, in-app video, or sales enablement file.
- Sync target: no sync, time-sync, phrase-sync, lip-sync, avatar lip-sync, or re-edited native cut.
- Review path: who signs off terminology, claims, performance, accessibility, and final delivery.

## Prepare source assets before translation

Do not start from an auto-transcript alone unless no better source exists. Create a source localization packet:

- Locked source video with timecode reference.
- Dialogue transcript with speaker IDs, timecodes, overlapping speech notes, offscreen/onscreen status, and inaudible flags.
- Intent notes per section: joke, warning, emotional beat, product claim, CTA, legal disclaimer, safety instruction, character relationship, or plot reveal.
- On-screen text, UI strings, graphics, lower thirds, captions, subtitles, title cards, end cards, supers, and legal slates.
- Pronunciation list for names, products, acronyms, places, invented terms, and domain terminology.
- Music/effects/dialogue stems when dubbing or re-mixing.
- Rights notes for talent, music, archive footage, user-generated clips, public figures, minors, and synthetic processing.

If the source is not locked, mark every downstream artifact as provisional. A changed source edit can invalidate subtitles, dubbing timing, localized graphics, and approvals.

## Translate for purpose, then adapt for time

Use a two-pass script process:

1. Meaning pass: translate all claims, instructions, story beats, emotional intent, terminology, names, numbers, safety/legal statements, and calls to action accurately for the target locale.
2. Adaptation pass: reshape lines for reading speed, phrase duration, mouth visibility, performer breath, market idiom, brand voice, and on-screen context.

Maintain a line table for dubbed work:

| Field | Purpose |
| --- | --- |
| Source timecode in/out | Preserve timing and make retakes traceable. |
| Speaker | Support casting, subtitle labels, and mix decisions. |
| Source line | Keep reviewers anchored. |
| Literal meaning | Prevent adaptation from drifting away from intent. |
| Localized performance line | The actual recorded/synthesized line. |
| Sync constraint | No sync, time-sync, phrase-sync, lip-sync, off-camera, or can recut. |
| Pronunciation | Names, product terms, acronyms, regional terms. |
| Reviewer notes | Brand/legal/cultural/accessibility notes and status. |

For subtitles, preserve meaning and readability rather than every source word. For captions, preserve audio information. For dubbing, preserve performance intent and timing while remaining natural in the target language.

## Build glossary and style controls

Create a project termbase before batch localization:

- Product names: translate, transliterate, or leave unchanged.
- UI strings: match shipped product UI exactly; do not invent local labels.
- Brand slogans and campaign lines: mark approved translations or "transcreate, do not literalize."
- Legal, safety, medical, financial, and compliance terms: require subject-matter approval.
- Units, currency, dates, phone numbers, addresses, measurements, honorifics, and formality/register.
- Competitor names, trademarks, cultural references, idioms, banned phrases, and sensitive words.
- Pronunciation: IPA, plain-language pronunciation, stress, and acceptable variants.

For multi-market campaigns, create a global glossary plus locale overrides. Do not force one Spanish/French/Arabic/etc. line across markets when usage, regulation, or audience expectation differs.

## Cast and direct voices

Match communicative function before surface similarity:

- Role: narrator, expert, customer, founder, instructor, character, child, elder, support agent, avatar, announcer.
- Performance: warm, urgent, credible, playful, restrained, premium, conversational, technical, reassuring, comedic.
- Vocal attributes: range, pacing, energy, articulation, age impression where relevant, accent/dialect, breathiness, authority, texture.
- Cultural fit: native or near-native target locale performance, idiom comfort, correct register and pronunciation.
- Continuity: recurring brand voice, character continuity, multi-episode consistency.

Do not cast by stereotype. If a brief asks for a protected-class-coded voice without a production reason, reframe toward performance, locale, and audience comprehension.

For synthetic voices:

- Confirm the voice license allows the intended language, territory, media, duration, paid usage, client, derivative works, and future edits.
- Do not clone or imitate a real person, client employee, performer, celebrity, public figure, private individual, child, deceased person, or previous actor unless explicit, informed, written permission covers the exact use.
- Keep a consent record, source of voice, model/provider, approved use, expiration, revocation terms, and disclosure requirements.
- Avoid "sound like [living actor/celebrity/creator]" directions. Use production descriptors instead.
- Disclose AI use where platform policy, client policy, law, or audience trust requires it. When unsure, escalate for policy/legal review.

## Direct sync and performance

For phrase-sync:

- Align sentence starts, ends, pauses, and emotional beats to the source more than individual mouth shapes.
- Rewrite instead of speeding target audio into an unnatural delivery.
- Use edit points, reaction shots, B-roll, screen recordings, or graphics to hide unavoidable expansion.
- Keep breaths and pauses believable; audiences hear unnatural compression even when they cannot read the language.

For lip-sync:

- Identify "hero sync" shots: close-ups, front-facing mouths, high-emotion lines, product claims delivered to camera, and avatar monologues.
- Mark "low sync" shots: off-camera, profile, wide shot, masked mouth, fast montage, B-roll, graphics, or cutaway.
- Adapt lines to fit visible mouth closures, phrase length, and emotional beat; do not sacrifice legal/medical/safety accuracy.
- Prioritize visible plosives and mouth closures at line ends and obvious close-ups, but do not over-optimize phonetics at the expense of natural speech.
- Use re

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [calesthio](https://github.com/calesthio)
- **Source:** [calesthio/generative-media-skills](https://github.com/calesthio/generative-media-skills)
- **License:** MIT
- **Homepage:** https://github.com/calesthio/OpenMontage

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-calesthio-generative-media-skills-localization-dubbing-production
- Seller: https://agentstack.voostack.com/s/calesthio
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
