AgentStack
SKILL verified MIT Self-run

Captions Media Accessibility

skill-calesthio-generative-media-skills-captions-media-accessibility · by calesthio

Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video recuts, training media, and localized content. Use when planning, authoring, reviewing, localizing, burning in, exporting, or QAing captions, subtitles, SDH, transcripts, audio description, flashing/motio…

No reviews yet
0 installs
4 views
0.0% view→install

Install

$ agentstack add skill-calesthio-generative-media-skills-captions-media-accessibility

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Captions Media Accessibility? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Captions and media accessibility direction

Treat captions and accessibility tracks as production assets, not polish. Build them from the script, audio mix, edit, localization plan, and platform delivery target. If the user asks for legal compliance, say what standards you are using and recommend review by qualified accessibility/legal counsel; do not promise legal compliance from a generated file alone.

Start by classifying the media

Before writing or exporting captions, identify:

  • Content type: prerecorded video with audio, live/near-live video, audio-only, video-only, or silent social cut.
  • Audience need: captions for deaf/hard-of-hearing viewers, translated subtitles for hearing viewers, SDH, transcript, descriptive transcript, audio description, sign language, or platform-default auto captions.
  • Delivery target: social burn-in, web player sidecar, broadcast/OTT, LMS/training portal, internal review, localized package, or archival master.
  • Accessibility risk: essential visual text, charts, screen recordings, multiple speakers, overlapping dialogue, music/lyrics, sound-driven story beats, flashing/strobe, fast kinetic typography, low contrast, small mobile screens, or platform UI overlays.
  • Final-source truth: the locked audio mix and final picture. Generated scripts, TTS drafts, and ASR transcripts are starting points, not the authority after edits.

Documented facts to preserve

Use these as factual constraints when relevant, citing the source in your output or handoff notes when the claim matters.

  • WCAG 2.2 requires captions for prerecorded audio in synchronized media at Level A (SC 1.2.2), captions for live synchronized audio at Level AA (SC 1.2.4), and audio description for prerecorded video at Level AA (SC 1.2.5). WCAG Level A allows either audio description or a media alternative for prerecorded synchronized video under SC 1.2.3. Source: W3C WCAG 2.2, verified 2026-07-10.
  • WCAG 2.2 flashing guidance says content must not flash more than three times in any one-second period unless below the general/red flash thresholds. Source: W3C Understanding SC 2.3.1, verified 2026-07-10.
  • WCAG contrast minimum for text is 4.5:1, with 3:1 for large-scale text. This is a web-text criterion; burned-in captions are video pixels, but use it as the minimum design target for caption readability unless a stricter platform/spec applies. Source: W3C WCAG 2.2 SC 1.4.3, verified 2026-07-10.
  • Section508.gov distinguishes closed captions, open captions, subtitles, and transcripts: closed captions can be turned on/off; open captions are permanent in the video; subtitles usually translate dialogue for hearing audiences and do not normally include speaker IDs or non-speech audio; transcripts are not time-coded. Source: Section508.gov synchronized media guidance, verified 2026-07-10.
  • U.S. FCC caption quality rules for covered television programming define quality around accuracy, synchronicity, completeness, and placement; captions should include relevant nonverbal information, be legible, and avoid blocking essential visual content. Source: 47 CFR Sec. 79.1 via Legal Information Institute/eCFR text, verified 2026-07-10. This skill does not determine whether a specific project is covered by FCC rules.
  • YouTube Help listed SRT and SBV/SubViewer as basic formats and WebVTT/TTML/SAMI/RealText as advanced formats; basic SRT/SBV files must be plain UTF-8 and YouTube notes that styling support is limited. Source: YouTube Help, verified 2026-07-10.
  • Vimeo Help stated that Vimeo supports SRT and WebVTT for captions/subtitles, recommends WebVTT, and requires UTF-8 encoding for special characters. Source: Vimeo Help, verified 2026-07-10.
  • WebVTT is a UTF-8 timed-text format used with media tracks; it begins with WEBVTT, contains timed cues, and can be used for captions, subtitles, descriptions, chapters, or metadata depending on player support. Sources: W3C/MDN and Library of Congress WebVTT documentation, verified 2026-07-10.
  • W3C WebVTT syntax requires UTF-8, a WEBVTT file signature, cue timings with period milliseconds, ordered cue start times, and end times greater than start times. Cue overlap is allowed by the format for captions/subtitles, but many caption delivery workflows still reject or flag overlaps. Source: W3C WebVTT Candidate Recommendation Draft, verified 2026-07-11.
  • SRT is a widely used but lightly standardized plain-text sidecar convention made of repeated blocks: numeric counter, start/end timing line, caption text, and a blank-line separator. Common delivery practice uses HH:MM:SS,mmm --> HH:MM:SS,mmm timing and UTF-8 for modern platform upload, even though legacy SRT files may use other encodings. Source: Library of Congress SRT format description plus platform guidance already cited here, verified 2026-07-11.
  • Netflix timed-text guidance is platform-specific, not universal law. Its public partner help guidance includes 2-line maximums, language-specific style guides, minimum/maximum event durations in general requirements, and U.S. English SDH conventions for bracketed speaker IDs/sound effects. Source: Netflix Partner Help Center, verified 2026-07-10.

Production heuristics to apply

These are craft defaults. Override them for a formal client style guide, broadcaster spec, language-specific subtitle convention, or platform validation error.

  • Prefer a sidecar caption file for accessibility whenever the platform supports it. Add burned-in/open captions only when the platform, ad buy, or user experience requires always-visible text.
  • If burning in captions for social, still keep an editable sidecar/transcript master. Burned-in captions cannot be toggled, restyled, searched, translated, or reliably used by assistive tech.
  • Use edited captions for prerecorded media. Auto captions are useful for draft alignment, but final work needs human/agent review against audio, speaker identity, punctuation, sound effects, proper nouns, and timing.
  • Keep cues readable: usually 1-2 lines, balanced line lengths, natural phrase breaks, and enough on-screen time to read without racing. If text is too dense, edit only nonessential redundancy while preserving meaning and key vocabulary.
  • Time cues to the audio they represent: appear near speech/sound onset, disappear near its end, avoid overlapping cues, and avoid making captions lead jokes, reveals, or safety-critical instructions before they are heard.
  • Reposition captions when they cover faces, lower-thirds, product UI, charts, subtitles already in picture, legal supers, or platform controls.
  • In generated video, reduce caption work by designing the script and edit for accessibility: fewer competing voices, readable on-screen text, visual information described in narration when possible, and pauses for audio description where needed.

Caption and SDH authoring

Create a caption master from these inputs:

  • locked video and audio;
  • final narration/dialogue script or transcript;
  • speaker list with names, titles, roles, pronunciations, and when names are first revealed;
  • glossary for product names, proper nouns, acronyms, technical terms, legal phrases, and invented/generated names;
  • music cue sheet and lyrics status;
  • on-screen text list, charts/data callouts, and any visual-only story beats;
  • target format(s), frame rate/timecode policy, language/locale tag, and style guide.

For each caption cue:

  • Include spoken words in order. Do not paraphrase unless readability/time constraints require compression and the chosen standard permits it.
  • Preserve meaning, tone, and essential vocabulary. Keep intentional slang, grammatical errors, repetitions, or hesitations when they matter to characterization, humor, safety, or instruction.
  • Use punctuation to clarify meaning, not to decorate.
  • Identify speakers when the viewer cannot reliably infer who is speaking from placement or picture. Use known names only after the content reveals them; otherwise use neutral descriptors such as [narrator], [offscreen voice], [woman 1], or client-approved labels.
  • Include non-speech audio when it affects meaning, mood, safety, story, or instruction: [alarm blares], [soft piano music], [audience laughs], [door unlocks]. Do not caption every incidental sound.
  • Caption lyrics when lyrics are audible and important, subject to rights/client instructions. If lyrics are not being captioned, identify the music when it affects the experience.
  • Avoid exposing information not available to hearing viewers at that moment. Do not name a mystery speaker, reveal a twist, or explain a hidden cause early.
  • Avoid visual descriptions inside captions unless they are representing sound or speaker identification. Visual access belongs in narration, integrated description, audio description, or descriptive transcript.

Line breaks:

  • Break at sentence, clause, or phrase boundaries.
  • Avoid separating articles from nouns, adjectives/modifiers from nouns, first and last names, auxiliary verbs from main verbs, prepositions from their objects, or a subject pronoun from its verb.
  • Do not place the end of one sentence and the beginning of another on the same line when a cleaner split is available.
  • Keep two-speaker cues rare; when unavoidable, make the speakers visually distinct with a consistent convention supported by the target format/platform.

Timing:

  • Align captions to final audio, not script timing.
  • Do not let captions persist far beyond speech unless needed for readability and not misleading.
  • Use gaps when they improve readability, but avoid unnecessary flicker from very short gaps.
  • If a speaker is interrupted, use punctuation that makes the interruption understandable.
  • For fast disclaimers, safety instructions, or legal text, push back on unreadable delivery instead of hiding the problem in tiny captions.

Burned-in/open caption design

Burn-in is typography over moving imagery. Design it like a legible UI:

  • Use high contrast against all representative frames. Add a semi-opaque box, shadow, stroke, or gradient when backgrounds vary.
  • Keep text large enough for the smallest expected screen and compression path. Preview at phone size, not just on a desktop monitor.
  • Leave margins for platform UI: captions, handles, reaction buttons, progress bars, lower-thirds, and safe areas differ by platform and orientation.
  • Avoid placing captions over faces, mouths in lip-sync/avatar work, products being demonstrated, charts, subtitles already present, and sign-language interpreters.
  • Keep animations restrained. Kinetic word-by-word captions can help social retention, but they are not a substitute for accessible captions if they omit words, move too quickly, lack contrast, or cannot be turned off.
  • Do not use color alone to distinguish speakers or meaning. Pair color with labels, placement, or other non-color cues.

Transcripts and descriptive transcripts

Provide a transcript when the user asks for one, when the media is audio-only, when training/search/review workflows benefit from it, or when accessibility requirements call for a time-based media alternative.

Basic transcript:

  • Include speech, speaker names, meaningful non-speech audio, and music/lyrics notes.
  • Clean obvious ASR errors, punctuation, and proper nouns.
  • Add headings or timestamps when useful for navigation.

Descriptive transcript:

  • Include the basic transcript plus essential visual information: on-screen text, charts, actions, scene changes, demonstrations, gestures, and visual-only jokes or instructions.
  • Use the same source of truth as captions and audio description.
  • Prefer concise descriptions that let a reader understand the content without seeing or hearing the video.

Audio description planning

Plan audio description before picture lock when possible. The cheapest and cleanest solution for many explainers, demos, and training videos is integrated description: write the main narration so it naturally says essential visual information.

Choose the method:

  • Integrated description: best for new explainers, training videos, screen demos, product walkthroughs, and educational media where narration can mention what matters visually.
  • Separate audio-description track: best when the original mix should remain unchanged and the player/platform supports an alternate described audio track.
  • Described-video alternate export: best when the platform does not support user-selectable audio description but can host a second version.
  • Timed text description file: possible where the player supports descriptions from text; verify playback support.
  • Descriptive transcript: useful for broad access and required in some WCAG contexts, but not always a replacement for synchronized audio description at the requested conformance level.

What to describe:

  • Essential actions, identities, scene changes, settings, charts, diagrams, screen operations, on-screen text, product states, gestures, and visual-only cause/effect.
  • Describe from general to specific when orienting a viewer.
  • Prioritize what is needed to follow, understand, and appreciate the content. Do not fill every pause.
  • Use present tense, active voice, objective wording, and vocabulary suited to the audience.
  • Avoid interpreting emotion or intent when a visible fact is enough. Prefer "Maya clenches her fists" over "Maya is furious" unless the emotion is otherwise stated or unmistakable and relevant.
  • Place description before or during the relevant visual moment when possible, not after the viewer needs it.
  • Avoid talking over dialogue or audio that is essential to comprehension unless necessary.

Flashing, motion, and sensory safety

Run a sensory-safety pass for generated motion graphics, glitch edits, lightning, alarms, camera flashes, strobes, rapid cuts, high-contrast pattern flicker, red flashes, and aggressive kinetic typography.

  • Avoid flashes above the WCAG three-flashes-per-second threshold unless a qualified tester confirms the flash is below the general/red flash thresholds.
  • Reduce or remove full-screen white/red flashes, alternating high-contrast frames, rapid inversion effects, and repeated strobe transitions.
  • Avoid making captions themselves flash, shake, spin, blur, or rapidly change position.
  • For motion-heavy deliverables, consider a reduced-motion variant or less intense edit when the audience or platform warrants it.
  • A warning is not a substitute for avoiding unsafe flashing when you can design it out.

Localization handoff

Do not treat localization as "translate the SRT." Prepare a package:

  • final video reference with burnt-in review timecode if helpful;
  • source transcript and captions;
  • speaker list, character names, pronouns if provided, titles/roles, and pronunciation notes;
  • glossary for product names, brand terms, acronyms, UI strings, legal phrasing, jokes, and do-not-translate terms;
  • explanation of audience, register/formality, locale, reading level, and platform;
  • notes for songs, lyrics, quoted material, censored/bleeped words, and foreign-language dialogue;
  • separate instructions for subtitles versus SDH in the target language;
  • screenshots or markers for on-screen text that needs translation, dubbing, replacement, or description;
  • required output formats, locale tags, and file naming.

Localization decisions:

  • Produce separate files per language/locale and per purpose: translated subtitles, same-language captions/SDH, forced narrative subtitles, audio-description scripts, and transcripts are distinct deliverables.
  • Let target-language reading speed and subtitle conventions govern line breaks, punctuation, and compression.
  • Preserve speaker identification and meaningful non-speech audio in SDH. Standard translated subtitles for hearing audiences may omit non-speech audio unless the client asks for SDH-style subtitles.
  • Do not translate source-language dialogue inside same-language captions unless the deliverable is explicitly interling

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.