Install
$ agentstack add skill-calesthio-generative-media-skills-avatar-spokesperson-production ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Avatar Spokesperson Production
Use this skill to plan, direct, review, and package synthetic presenter videos where a humanlike avatar or digital spokesperson speaks to the viewer. It is provider-independent: adapt the workflow to the avatar, TTS, dubbing, lip-sync, compositing, or platform tools available in the current environment.
Do not treat avatar production as only a prompt-writing task. The core production risks are identity rights, audience transparency, claims accuracy, performance credibility, brand trust, localization, and QA.
This skill is not legal advice. When rights, advertising claims, union/talent terms, regulated industries, minors, political/civic content, health/financial claims, public figures, or deceased personalities are involved, escalate to the client, counsel, brand owner, performer/voice owner, or compliance reviewer before generation or publication.
Non-negotiable production boundaries
Separate documented facts, project requirements, and production heuristics:
- Documented fact: cite a source or project artifact.
- Project requirement: comes from the client brief, contract, brand guide, legal review, consent form, platform spec, or tool documentation.
- Production heuristic: a practical quality rule that improves outcomes but is not a legal or technical requirement.
Before generation, confirm:
- The avatar and voice are licensed or consented for this use, audience, territory, duration, channels, languages, and editing scope.
- The video will not imply a real person, employee, expert, customer, celebrity, government official, medical/legal/financial professional, or organization endorsed a claim unless that is true and approved.
- Required AI/synthetic-media disclosures and platform labels are planned.
- Objective claims have substantiation and high-risk claims have the right approval.
- The script, pronunciation list, performance direction, wardrobe/background, captions, and delivery variants are production-ready.
First pass: classify the avatar job
Capture these fields before selecting tools or writing prompts:
project:
goal: "What business or learning outcome must the avatar achieve?"
audience: "Who is watching, where, and what do they already believe?"
format: "corporate explainer | training | sales | support | localization | announcement | compliance | social"
channels: ["landing page", "YouTube", "TikTok", "LinkedIn", "LMS", "internal portal"]
duration_targets: ["primary 60s", "vertical 30s cutdown", "6s hook"]
risk_level: "low | moderate | high"
persona:
role: "host | instructor | support agent | product manager | executive proxy | narrator"
identity_type: "licensed real person | synthetic non-identifiable avatar | employee likeness | public figure likeness"
voice_type: "licensed human voice | synthetic voice | cloned voice | recorded actor | localized dub"
consent_artifacts: ["contract", "release", "usage limits", "revocation/expiry terms"]
content:
script_source: "new | client-provided | translated | regulated/compliance"
claims_inventory: ["all objective claims, endorsements, statistics, comparisons"]
forbidden_implications: ["what the avatar must not appear to say, endorse, or represent"]
localization:
languages: ["en-US"]
pronunciation_list: ["brand names", "people names", "acronyms", "technical terms"]
cultural_notes: ["gestures", "formality", "dress", "examples to avoid"]
approvals:
required_reviewers: ["brand", "legal", "compliance", "talent/performer", "regional reviewer"]
Set risk_level to high if any of the following apply: public figure or employee likeness; cloned voice; medical, legal, financial, political, employment, housing, education, or safety claims; children or vulnerable audiences; synthetic expert/customer testimonials; regulated training; crisis communications; international release; or a platform where synthetic-media policy is material to distribution.
Rights, consent, and likeness review
Treat avatar identity and voice as separate rights surfaces. A realistic face, body, voice, name, title, uniform, accent, and mannerism can each create audience assumptions.
Minimum consent packet for a real or cloned person:
- Written permission for avatar/voice capture, model creation if applicable, generation, editing, storage, and distribution.
- Specific permitted uses: product, campaign, client, channels, territories, languages, duration, paid media, organic social, internal training, and derivatives.
- Explicit prohibited uses: political persuasion, adult content, medical/financial advice, competitive brands, crisis statements, new scripts after expiry, model training, face/voice transfer, or unsupervised self-service.
- Compensation, revocation, renewal, exclusivity, takedown, and audit terms.
- Approval workflow for scripts, likeness, voice, translated performances, and final cuts.
- Security requirements for source recordings, voiceprints, face scans, prompts, project files, and generated takes.
SAG-AFTRA summarizes its AI guardrails as consent, fair compensation, and control over performances. NAVA recommends performer contracts cover consent, use limits, opt-out or term limits, payment, exclusivity, and safe storage/tracking of voice, likeness, performance, and resulting products. Use these as production governance principles even when the project is non-union, then defer to the actual contract and counsel for binding terms.
For a synthetic non-identifiable avatar, still check whether the avatar resembles a real person, implies a protected role, or uses a voice trained on identifiable speakers. Do not describe a synthetic avatar as an actual employee, doctor, customer, founder, or public official unless the identity and approval are real.
Disclosure and labeling plan
Plan disclosure as part of the creative, not as an afterthought. The disclosure should be clear enough that an ordinary viewer understands that the presenter is synthetic or AI-generated when that fact is material to trust, identity, or platform policy.
Good disclosure locations:
- Opening or lower-third text for high-risk or testimonial-like content: "AI-generated avatar spokesperson."
- End card or description for lower-risk internal explainers: "Presenter video created using an AI avatar and synthetic voice."
- Platform-specific upload toggles or AI labels.
- Campaign trafficking notes so paid-media, social, and web teams do not strip the disclosure.
Avoid disclosures that are tiny, fleeting, hidden behind a click, contradicted by the script, or placed only after a misleading impression has already formed.
For advertising and endorsements, the FTC endorsement guides address endorsements and material connections under Section 5, and FTC advertising substantiation policy requires a reasonable basis for objective claims. If an avatar states a product claim, expert opinion, customer experience, ranking, statistic, savings number, health/safety outcome, or comparative superiority claim, the client must provide support before publication. Do not let the realism of an avatar create a fake testimonial or expert endorsement.
Platform and legal rules change. Re-check platform synthetic-media requirements at production time, especially for YouTube, TikTok, Instagram/Facebook, paid ads, app stores, LMS platforms, and regional variants. As verified on 2026-07-11: YouTube requires disclosure for realistic altered/synthetic content viewers could mistake for a real person, place, scene, or event; TikTok requires labels for AI-generated content containing realistic images, audio, or video; Meta describes "AI info" labels based on industry-standard signals and self-disclosure for AI-generated or AI-modified video, audio, and image content. Treat these facts as volatile.
Script and persona brief
Write for an avatar like a camera-facing presenter, not a generic narrator. The script should be easy to perform in one breath group at a time and should not require facial nuance the avatar cannot deliver.
Create a persona brief:
persona_brief:
speaker_role: "Calm onboarding coach for new enterprise users"
audience_relationship: "helpful peer, not executive authority"
credibility_basis: "company-approved training host; not a real customer or lawyer"
warmth: "medium-high"
formality: "plain professional"
pace: "145 words per minute target; slower for compliance terms"
energy: "confident but not salesy"
eye_line: "direct-to-camera for explanations; glance down only for quoted checklist moments"
gestures: "small open-hand emphasis on transitions; no pointing at viewer"
wardrobe: "solid navy blazer, no lab coat or official uniform"
background: "brand-neutral office gradient, no fake newsroom/clinic/courtroom"
forbidden_read: "must not appear to be a real employee giving legal advice"
Script rules:
- Put the key promise in the first 3-5 seconds for social and product announcements.
- Keep sentences short. Prefer one idea per sentence.
- Use contractions only if appropriate to brand voice and localization.
- Avoid tongue-twisters, repeated sibilants, awkward alliteration, and dense clauses near visual transitions.
- Spell out numbers and units when TTS may misread them: "twenty-four seven," "SOC two," "A P I."
- Mark pronunciation:
OpenMontage (OH-pun mon-TAHZH),GDPR (G D P R),SaaS (sass). - Mark pauses and emphasis sparingly:
[beat],[slower],[smile],[firm]. - Avoid saying the avatar has personal experience unless the experience belongs to the actual licensed person and is approved.
For compliance or training modules, separate "must-say" approved copy from optional connective tissue. Do not paraphrase regulated language unless the compliance reviewer approves.
Casting and representation
Casting should support comprehension and trust without tokenism, stereotype, deception, or false authority.
Check:
- Does the avatar's apparent age, role, attire, accent, and setting imply credentials or lived experience?
- Would the same script be appropriate if spoken by a real person with this identity presentation?
- Is the avatar being used to simulate a customer, patient, student, doctor, lawyer, public official, or worker? If yes, require review.
- Are race, gender, disability, accent, nationality, body type, or dress being used as shorthand for trust, expertise, service labor, or exoticism?
- Does localization adapt persona, formality, gestures, examples, and pronunciation rather than merely translating words?
Prefer role clarity over realism. "Synthetic product guide" is often safer and more honest than "AI customer testimonial."
Performance direction
Avatar providers vary in how much control they expose. Even when controls are limited, write direction in production language so the prompt, script markup, and review are aligned.
Direct:
- Eye-line: direct, slightly off-camera interview, teleprompter, or screen-pointing.
- Gesture density: none, low, medium, high. Corporate explainers usually need low-to-medium gestures; compliance videos usually need low gestures.
- Facial expression: neutral, warm, reassuring, serious, celebratory. Avoid rapid emotional shifts.
- Head movement: stable for authority, mild nods for support, no bobbing.
- Pacing: words per minute, pauses after section headers, slower pronunciation for names, acronyms, and legal language.
- Shot type: bust, waist-up, full-body, seated, standing, picture-in-picture, screen-side presenter.
- Interaction with graphics: indicate whether the avatar should point, glance, or leave space for overlays. If the provider cannot reliably coordinate gaze/gesture with graphics, use post-production graphics instead of generated pointing.
Break long scripts into takes. Generate and review 10-20 second segments for high-risk work before committing to a full-length render. For localization, generate a short proof clip per language that includes the hardest names, numbers, and claims.
Wardrobe, background, and brand fit
Make the avatar fit the brand without implying unauthorized affiliation or credentials.
Wardrobe:
- Use solid colors and simple patterns to reduce rendering artifacts.
- Avoid uniforms, badges, lab coats, judicial robes, military/police-like attire, or branded logos unless authorized and needed.
- Avoid jewelry or fine patterns that shimmer or deform during lip-sync.
Background:
- Choose a setting that matches the promise: product UI, support desk, training room, neutral studio, event announcement, or brand gradient.
- Avoid fake newsrooms, clinics, courtrooms, government offices, exchange trading floors, or classrooms when they imply authority the avatar does not have.
- Leave safe zones for captions, lower thirds, platform UI, and crop variants.
Brand:
- Apply brand colors to background, lower thirds, slides, and supers rather than forcing the avatar to wear logo-heavy clothing.
- Keep contrast high enough for captions and titles.
- Ensure product screenshots, UI states, prices, and legal footers are current and approved.
Claims, endorsements, and sensitive-content review
Create a claims inventory before production:
claims_inventory:
- script_line: "Cut onboarding time by 40%."
claim_type: "objective performance claim"
substantiation: "client study Q2 2026, approved by marketing/legal"
risk: "moderate"
reviewer: "legal"
- script_line: "I recommend this to every sales team."
claim_type: "endorsement/testimonial"
substantiation: "none"
action: "rewrite; avatar cannot imply personal recommendation"
Escalate if the avatar:
- Says "I tried," "my patients," "my clients," "as your doctor/lawyer/advisor," "we guarantee," or "officially approved" without proof.
- Mentions prices, discounts, availability, earnings, savings, performance, safety, health outcomes, environmental benefits, or comparative superiority.
- Looks or sounds like a real person who did not approve the exact use.
- Speaks about political, civic, emergency, medical, legal, financial, employment, housing, education, or insurance decisions.
- Is used in paid advertising, influencer-style content, testimonials, or native ads.
Rewrite synthetic endorsements into transparent brand narration:
- Risky: "I use Acme every day, and it saved me four hours."
- Safer: "Acme's customer study found an average four-hour weekly time savings among surveyed teams."
- Risky: "As your financial advisor, I recommend this plan."
- Safer: "This explainer is not financial advice. Review the plan details with a qualified advisor."
Localization and pronunciation
Localization quality is a trust issue. A visually polished avatar with wrong pronunciation or culturally odd gestures feels deceptive or careless.
For each language or market:
- Use an approved translation or transcreation, not raw machine translation for customer-facing or regulated work.
- Create a pronunciation table for brand names, product names, people, places, acronyms, numbers, and domain terms.
- Confirm locale-specific formality, honorifics, units, currency, date formats, privacy terms, and examples.
- Check gesture norms. A gesture that reads friendly in one market can read dismissive or confusing in another.
- Decide whether to use the same avatar across markets or localize presenter appearance/voice. Do not assume one avatar is culturally neutral.
- Generate a proof segment that contains the hardest words and a representative emotional passage.
- Have a native or market-qualified reviewer approve audio, captions, on-screen text, and any mouth-shape oddities.
Keep source text, translated script, back-translation notes, pronunciation dictionary, voice/avatar choice, and reviewer signoff in the delivery record.
Lip-sync, audio, and visual QA
Review avatar video at normal speed, half speed, and with audio-only playback.
Lip-sync and face:
- Mouth closures match
p,b,m; teeth/tongue do not flicker
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: calesthio
- Source: calesthio/generative-media-skills
- License: MIT
- Homepage: https://github.com/calesthio/OpenMontage
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.