Install
$ agentstack add skill-calesthio-generative-media-skills-explainer-video-production ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Explainer video production
Produce an explainer only after defining what the audience should understand, believe, decide, or do differently after watching. Treat the video as a learning and decision artifact, not as a sequence of pretty scenes.
This skill is provider-independent. Use the available generation, editing, composition, captioning, and review tools in the host environment, but keep the explainer logic, factual discipline, accessibility, and QA standards intact across providers.
Non-negotiables
- Start from audience, prior knowledge, and a measurable learning or decision objective.
- Separate documented facts, source-backed inferences, production heuristics, and creative inventions.
- Maintain a claim log for factual, commercial, comparative, health, safety, legal, financial, scientific, and public-service claims.
- Escalate high-risk claims to the client, subject-matter expert, legal/compliance reviewer, medical reviewer, financial reviewer, or policy owner. Do not provide legal, medical, or financial advice.
- Build accessibility into script, storyboard, captions, audio, data visuals, and delivery variants. Retrofitting is weaker and more expensive.
- Re-check volatile provider/platform facts at production time: social aspect ratios, maximum lengths, safe areas, caption behavior, model inputs, model limits, pricing, regional availability, disclosure rules, and rights/licensing terms.
- Never ask an image or video generation model to invent evidence, render precise charts, preserve exact legal/medical/scientific statements, or create readable fine text. Use deterministic layout/composition for facts, citations, captions, tables, UI text, labels, and charts.
Intake: define the learning contract
Before writing or generating, answer:
- Audience: Who is this for, what do they already know, what do they misunderstand, what language/register do they use, and what accessibility/localization needs are known?
- Objective: What single sentence should the viewer be able to say, do, or decide after watching?
- Use case: education, product, onboarding, training, public-service, fundraising, policy, sales enablement, social awareness, internal change management, or technical concept.
- Success evidence: quiz answer, demo completion, reduced support question, sign-up, policy comprehension, behavior change, stakeholder approval, or share/save.
- Constraints: duration, platform, brand, required sources, must-include/must-avoid claims, restricted imagery, voice, captions, languages, review approvers, budget, and tool availability.
- Risk tier:
- Low: general concept, internal orientation, non-sensitive creative explanation.
- Medium: product benefits, nonprofit/public-service behavior guidance, data claims, employment/training requirements.
- High: health, safety, legal, financial, regulated products, children, political persuasion, crisis guidance, comparative advertising, public statistics with policy implications.
Write the objective as an observable outcome, not a vague topic:
- Weak: "Explain carbon offsets."
- Strong: "After 90 seconds, a first-time buyer can tell the difference between reducing emissions and buying offsets, and knows to look for project verification before trusting a claim."
Research and citation discipline
Create a claim matrix before scripting. The matrix prevents persuasive polish from outrunning the evidence.
| Field | Use | |---|---| | Claim | The exact statement or implication the video may make. | | Claim type | Definition, statistic, causal, comparative, product performance, testimonial, safety, behavior recommendation, forecast, opinion, metaphor, or CTA. | | Source | Primary source preferred; include title, URL, publication/update date, and access date. | | Evidence strength | Official/primary, peer-reviewed, technical report, reproducible test, reputable secondary, client-provided, anecdotal, or creative. | | On-screen handling | Citation, source line, voice-only, legal/disclosure copy, excluded from final, or expert-review required. | | Risk | Low/medium/high and why. | | Owner | Agent, client, SME, legal/compliance, medical, financial, accessibility, localization. |
Use primary sources where possible: official documentation, standards, regulator guidance, government statistics, peer-reviewed research, product docs, client-approved evidence, or directly inspected source materials. Use secondary sources only when primary sources are unavailable or for background, and label them as such.
For commercial or product explainers, objective performance claims need substantiation. FTC business guidance says companies must support advertising claims with solid proof, especially health-related claims; the FTC Endorsement Guides also require material connections in endorsements to be disclosed clearly and conspicuously. Treat these as compliance triggers, not as copywriting suggestions. Re-check current FTC guidance and applicable jurisdiction at production time.
For AI-generated media and stock/media assets, record provenance: model/provider, prompt, seed if available, source asset license, release/consent status, edit history, and whether material is AI-generated. U.S. Copyright Office guidance says copyright registration of works containing AI-generated material turns on human authorship and disclosure of more-than-de-minimis AI-generated content; rights rules are volatile and jurisdiction-dependent, so escalate registration, ownership, likeness, and training-data questions to legal counsel.
Decompose the concept before writing
Break the idea into explainable atoms:
- Terms: vocabulary the audience needs before the explanation works.
- Parts: components, actors, inputs, outputs, constraints.
- Mechanism: what changes, in what order, and why.
- Contrast: what this is not; common misconception or false alternative.
- Example: concrete case, scenario, or mini-demo.
- Consequence: why it matters; what happens if ignored.
- Action: what the viewer should do next.
Choose the minimum set needed for the objective. If the concept has too many atoms for the duration, narrow the promise instead of compressing everything.
Useful decomposition patterns:
- Definition explainer: term -> contrast -> example -> use.
- Mechanism explainer: input -> transformation -> output -> feedback loop.
- Product explainer: problem -> current workaround -> product mechanism -> proof -> next step.
- Public-service explainer: situation -> risk -> recommended action -> exception -> where to get help.
- Training explainer: task goal -> prerequisites -> steps -> check -> recovery path.
- Data explainer: question -> dataset/source -> pattern -> caveat -> implication.
Script architecture
Use the structure that fits the objective; do not force every explainer into the same funnel.
Core structure for most 60-180 second explainers:
- Hook: name the tension or practical question.
- Promise: tell viewers what they will understand or be able to do.
- Map: preview the two to four pieces of the explanation.
- Build: explain one idea per beat, each beat resolving one question.
- Example/demo: show the concept working in a concrete case.
- Caveat or boundary: prevent overclaiming and address a common misconception.
- Action: next step, summary, practice task, or decision prompt.
Short social variant, 15-45 seconds:
- 0-3s: pattern interrupt or specific problem.
- 3-8s: promise and stakes.
- Middle: one mechanism or one before/after, not a full curriculum.
- End: memorable line, CTA, or "save this" recap.
Training/onboarding variant:
- State the task and success condition.
- Demonstrate steps in order.
- Add checkpoints and error recovery.
- End with where to find help or how to verify completion.
Public-service/nonprofit variant:
- Use plain language and avoid shame.
- Make the recommended action concrete.
- Show who the recommendation applies to and who needs different guidance.
- Cite official sources and escalate health/safety/legal claims.
Write narration for comprehension
Narration should sound like a capable guide, not a whitepaper read aloud.
- Use everyday words unless a technical term is necessary; define necessary terms before relying on them.
- Put one idea in each sentence.
- Keep clauses short enough for voiceover and captions.
- Use signposting: "First," "The key difference," "Here is the catch," "Now watch what changes."
- Repeat critical terms consistently; do not rotate synonyms for concepts the viewer is still learning.
- Put numbers in context: compare, convert, show scale, or say why the number matters.
- Read the script aloud. If the voice trips, the viewer will too.
For scripted voiceover, plan roughly 130-160 spoken English words per minute for calm explainers, slower for technical, translated, child-facing, or accessibility-sensitive videos. Treat this as a heuristic; actual pacing depends on language, speaker, audience, visuals, and pause needs.
Use analogies without lying
Analogies help when they map structure, not just mood.
Before using an analogy, write:
- What maps: "A is like B because both..."
- What does not map: "Unlike B, A does not..."
- Where to stop: the exact point after which the analogy becomes misleading.
Example: "A heat pump is like moving water uphill with a pump: you spend energy moving heat rather than creating heat. But unlike water, heat naturally flows from warmer to cooler places, so the device uses a refrigerant cycle to move it the other way."
Avoid analogies in high-risk domains if the simplification could change behavior, dose, safety, eligibility, legal interpretation, or financial decision-making. Use plain causal explanation instead.
Storyboard as information design
Storyboard every scene with these columns:
| Column | Required content | |---|---| | Time | Start/end, duration, and pacing intention. | | Learning beat | What the viewer learns or can now do. | | Narration | Exact voiceover or dialogue. | | Visual | Diagram, character action, screen capture, product UI, data chart, live footage, icon system, or generated media shot. | | Motion | What changes on screen and why that motion helps understanding. | | Text/caption | On-screen labels, source line, disclosure, lower third, or no text. | | Evidence | Claim IDs from the claim matrix. | | Accessibility | Caption note, audio-description need, contrast risk, no-color-only encoding, flashing/motion risk, transcript note. | | Asset/prompt notes | What must be deterministic vs. what may be generated. |
If a scene has no learning beat, cut it or convert it into a transition lasting only as long as needed.
Visual grammar for explainer scenes
Match the visual form to the cognitive job:
- Definition: term card, labeled object, side-by-side "is/is not."
- Process: flow, timeline, numbered sequence, conveyor, state machine.
- Cause/effect: before/after, causal chain, feedback loop, split screen.
- System: map of actors, inputs, outputs, dependencies, boundaries.
- Scale: familiar comparison, proportional bars, nested containers, map inset.
- Data: chart with one takeaway, highlighted trend, source line, caveat.
- Product: real UI or product surface, problem context, feature in use, result state.
- Training: cursor path, highlighted control, checklist, success/error state.
- Social/emotional: human scenario, testimonial with disclosure, character metaphor.
Keep visual language consistent: one color for "problem," one for "solution," one for "warning," one for "proof." Do not switch meanings mid-video.
Avoid decorative literalism. A cloud icon behind every mention of "cloud" rarely explains anything. Use visuals to show relationships, sequence, magnitude, change, and exceptions.
Diagrams, charts, and data callouts
Use charts and diagrams only when they answer a question better than narration.
Data and diagram rules:
- State the takeaway in words before or with the chart.
- Use the simplest chart that preserves the truth. Prefer bar charts for category comparisons, line charts for trends over time, timelines for sequence, and annotated diagrams for mechanisms.
- Do not use pie/donut charts for many categories, small differences, or precise comparisons.
- Label directly where possible; reduce legend hunting.
- Do not encode meaning by color alone. Add labels, position, shape, pattern, or text.
- For web or interactive delivery, provide accessible alternatives. W3C WAI treats graphs, charts, flow charts, diagrams, and maps as complex images when they contain substantial information, requiring short and long text alternatives. USWDS guidance notes that SVG chart content can be hard for screen readers and recommends screen-reader-accessible tables plus plain-text trend/statistical summaries, while warning that tables alone may not be sufficient for complex datasets.
- In video-only delivery, make the narration and captions carry the key chart takeaway, and provide a transcript or companion notes when the chart is essential.
- Generate charts deterministically from verified data. Do not ask a generative image model to create a factual chart.
Motion direction
Motion should explain change.
Use motion to:
- reveal sequence;
- connect cause to effect;
- show state transitions;
- compare before/after;
- guide attention;
- pace cognitive load;
- create emotional emphasis after the viewer understands the point.
Avoid motion that competes with the explanation:
- constant background motion under dense text;
- rapid camera moves during new information;
- simultaneous changes in multiple regions;
- transitions that imply causality where none exists;
- fake UI behavior that differs from the product;
- flashing more than three times in one second or flashes near known thresholds; WCAG includes a "three flashes or below threshold" requirement for web content.
Use pauses. A half-second hold after a key reveal often teaches more than another flourish.
Narration, voice, music, and sound
Choose voice by trust relationship:
- Teacher/mentor: calm, precise, warm.
- Product guide: confident, practical, not hype-heavy.
- Public-service: respectful, plain, non-shaming.
- Training: steady, directive, enough pauses for task execution.
- Youth/social: energetic but still intelligible and caption-friendly.
Prepare a pronunciation sheet for names, acronyms, medication/technical terms, non-English words, and brand terms. For TTS, test a short sample before batch generation.
Mix for speech intelligibility:
- Narration must remain dominant.
- Music should support pace and emotion, not fill every gap.
- Sound effects should mark state changes or actions; avoid cartoon clutter unless the style calls for it.
- If background audio runs longer than a few seconds under speech, keep it low and non-distracting. WCAG includes guidance on low/no background audio at AAA, and WAI media guidance emphasizes low background audio as part of accessible examples.
Captions, transcripts, and accessibility
Design accessibility from the script stage.
Minimum production targets:
- Captions for all speech and meaningful non-speech audio. WCAG 2.2 includes prerecorded captions at Level A and live captions at Level AA for synchronized media.
- Descriptive transcript for meaningful video where feasible; W3C WAI recommends transcripts that include speech, meaningful non-speech audio, and ideally visual information.
- Audio description or integrated description when key meaning is visual-only. WAI recommends planning integrated description before filming/scripting because it is easier and often better for accessibility.
- Sufficient contrast for text and essential graphical objects. WCAG 2.2 includes text contrast and non-text contrast requirements; use them as a floor, not a ceiling.
- No information conveyed only by color, sound, motion, or spatial position.
- Avoid dense text. If text must be read, leave enough on-screen time for the intended audience and
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: calesthio
- Source: calesthio/generative-media-skills
- License: MIT
- Homepage: https://github.com/calesthio/OpenMontage
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.