Install
$ agentstack add skill-calesthio-generative-media-skills-generated-media-qa ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Generated Media QA
Treat QA as a release decision, not a vibe check. Judge the deliverable against the approved brief, platform specifications, legal/safety constraints, and the audience context. Record enough evidence that another agent or producer can reproduce the decision.
Keep three evidence lanes separate
Documented facts are requirements from the brief, platform specs, legal/policy guidance, delivery standards, accessibility standards, or provider documentation. Cite them or name the source and verification date.
Empirical observations are what you directly measured or inspected in the asset: frame size, duration, loudness, sync offset, OCR output, transcript mismatch, visual artifact, metadata, or a timestamped defect.
Production heuristics are professional judgments used when no explicit spec exists: whether a hand artifact is audience-visible, whether a product packshot feels trustworthy, whether an accent is intelligible for the target market, whether a social caption is too fast for mobile. Label these as heuristics and avoid pretending they are universal standards.
Intake before review
Do not start with random artifact hunting. Build the acceptance frame first.
Collect:
- Approved brief, prompt, storyboard, script, shot list, edit decision list, brand rules, product facts, target audience, target platform, duration, aspect ratio, language/locale, and known compromises.
- Delivery spec: file format, codec, resolution, frame rate, color space/HDR, audio channels, loudness target, captions format, thumbnail and metadata requirements.
- Source inventory: generated assets, human-shot assets, licensed stock, user-provided media, logos, fonts, music, SFX, voices, product claims, model releases, likeness/voice consent, and usage rights.
- Generation metadata when available: provider, model, model version/date, prompt, negative prompt, reference images, seed, dimensions, duration, fps, voice ID, language, post tools, upscalers, edit tools, and safety filters.
- Risk context: ads, health/finance/legal claims, political content, public figures, minors, regulated products, synthetic endorsements, realistic news-like scenes, localization, accessibility obligations, and platform disclosure requirements.
If the brief is missing, write a provisional QA basis such as: "Reviewing against supplied asset, target platform: Instagram Reels, inferred goal: 15 s product teaser. Product-claim and legal review cannot pass until facts and rights are supplied."
First pass: acceptance matrix
Create a small matrix before deep inspection:
| Area | Pass evidence | Typical reject evidence | |---|---|---| | Brief conformance | Message, product, audience, tone, duration, format, and call to action match the approved brief | Wrong product, missing CTA, wrong locale, off-brand tone, materially changed promise | | Technical delivery | File opens, spec matches, no corruption, correct duration/fps/aspect, audio/captions present as required | Wrong aspect, corrupt frames, missing audio, bad encode, unsupported caption file | | Generated-media realism | No audience-visible anatomy, physics, identity, text, logo, continuity, or temporal defects | Warped hands/faces, unreadable text, drifting identity, morphing product, flicker, impossible motion | | Audio and speech | Intelligible, synced, no clipping/noise/pops, correct language/voice, loudness meets target | Dialogue buried, lip sync off, clipping, wrong voice, mistranslation | | Accessibility | Captions, transcript, alt text, audio description, safe flashing, and player affordances meet the delivery context | Missing captions where required, unreadable captions, unsafe flashes, visual-only information with no alternative | | Rights and provenance | Source rights, consent, metadata, and AI disclosure are documented | Unlicensed source, unapproved likeness/voice, missing synthetic disclosure, unverifiable origin | | Safety/policy | No disallowed deception, unsafe advice, discriminatory content, or platform-prohibited claims | Misleading synthetic person/event, unsupported health claim, unsafe instructions, regulated-product violation |
Then inspect. Do not let a single attractive frame override a failed acceptance criterion.
Technical file QA
Use automated checks for measurable properties and human review for perceptual defects. Automated file QC catches many delivery failures, but it cannot decide whether a generated product label is semantically wrong or whether a synthetic spokesperson feels deceptive.
Check:
- File integrity: opens in at least two players/viewers; no truncated tail, decode errors, missing streams, alpha surprises, or silent channels.
- Container and streams: expected format, codec, profile, resolution, pixel aspect ratio, frame rate, duration, color primaries/transfer/matrix, bit depth, bitrate, audio sample rate, channel layout, and captions/subtitle streams.
- Visual defects: black frames, freeze frames, dropped/duplicated frames, banding, macroblocking, moire, aliasing, unintended borders, watermark remnants, bad mattes, poor keying, crop errors, unsafe title area, unreadable small text, and compression damage.
- Audio defects: clipping, inter-sample peaks, noise floor, hum, clicks, pops, gating artifacts, reverb wash, harsh sibilance, phase cancellation, channel imbalance, missing stems, and abrupt edits.
- Versioning: filename, slate, burned-in timecode, metadata, export preset, and revision number match the delivery tracker.
For broadcast or premium delivery, use the client/platform spec as the authority. SMPTE describes IMF as a file-based media format for storing and delivering audiovisual masters across versions and territories. DPP/AS-11-style workflows use explicit technical and editorial metadata. These are not default requirements for a TikTok draft, but they are useful models for disciplined delivery tracking.
Visual and motion QA for generated images/video
Generated media needs targeted inspection beyond normal color and compression checks.
Inspect still frames at 100% and at the intended viewing size. Scrub video slowly, then watch once without pausing on the target device class. Many AI defects are only obvious during motion, and many "defects" disappear at normal mobile viewing size.
High-priority checks:
- Anatomy and faces: hands, fingers, teeth, eyes, ears, joints, limb count, asymmetric pupils, skin texture, hairline shimmer, identity drift, age mismatch, and face-morphing between frames.
- Text and symbols: product labels, UI text, subtitles burned into imagery, signs, legal copy, prices, dates, numbers, QR codes, logos, trademark shapes, and brand color placement. Use OCR when possible, but manually verify brand-critical text.
- Product truth: package geometry, materials, scale, ports/buttons, SKU variant, colorway, dosage/amount, included accessories, ingredients, claims, safety warnings, and compatibility details.
- Temporal continuity: subject identity, wardrobe, prop positions, lighting direction, shadows, reflections, weather, time of day, screen contents, labels, and geography across shots.
- Motion plausibility: foot sliding, floating objects, rubber limbs, non-causal physics, warped vehicles, rolling-shutter hallucinations, facial/lip deformation, camera path discontinuity, interpolation smear, and background crawling.
- Compositing and mixed-source edits: matte edges, grain/noise mismatch, perspective, lens blur, color temperature, contact shadows, reflection consistency, scale, eye lines, room tone, and source footage/generative insert seams.
- Avatar and lip-sync: mouth closure on bilabials, jaw timing, head/neck motion, eye blinks, gaze, breathing, hand gestures, language/accent match, uncanny stillness, and whether the avatar is disclosed/approved.
Production heuristic: accept small artifacts only when they are not visible at intended size, do not affect the subject/product/message, and will not become a meme or trust-breaker. Reject or revise any defect on a face, hand, product, logo, legal copy, price, medical/financial statement, or safety instruction.
Audio, loudness, and intelligibility
Use the delivery spec first. If none exists, choose a target appropriate to the platform and document it as a production heuristic.
Documented facts:
- ITU-R BS.1770 specifies algorithms for measuring programme loudness and true-peak audio level.
- EBU R 128 recommends an average programme loudness of -23 LUFS and the use of Loudness Range and Maximum True Peak descriptors for audio signals.
- ATSC A/85 provides methods for measuring and controlling loudness for digital television and is referenced by the FCC CALM Act context for TV commercials in the United States.
- Netflix partner guidance for its own deliveries tells reviewers to use meters implementing ITU-R BS.1770 variants and describes how/when to flag loudness issues. Treat Netflix requirements as Netflix-specific, not universal.
Review:
- Dialogue intelligibility: listen on headphones, laptop speakers, and a phone speaker when the target is social/mobile. Flag when music/SFX mask key words.
- Sync: check lip sync at start, middle, and end; check that translated dubs preserve turn-taking and emotional timing. For avatars, inspect consonant closure, pauses, breaths, and facial expression alignment.
- Loudness: measure integrated loudness, true peak, and loudness range. Record the tool and target. Do not "normalize by ear" for a final delivery.
- Mix balance: narration, dialogue, SFX, music, ambience, and captions should support the message. Lower music under legal disclaimers, product claims, and instructions.
- Localization: verify numbers, units, names, pronunciations, cultural references, idioms, and reading pace in the target language. Back-translation can catch gross errors but is not a substitute for native review on high-stakes content.
Captions, accessibility, and viewer safety
Accessibility is part of QA, not a last-minute export option.
Documented facts:
- WCAG 2.2 requires captions for prerecorded audio content in synchronized media at Success Criterion 1.2.2 and audio description or a media alternative at 1.2.3; Level AA includes audio description at 1.2.5 for prerecorded synchronized media.
- WCAG 2.2 Success Criterion 2.3.1 says content must not flash more than three times in any one-second period unless below the general and red flash thresholds.
- Netflix timed-text guidance is a platform-specific professional reference for subtitle timing, duration, forced narratives, and style. Use the actual client/platform guide when delivering elsewhere.
Check:
- Captions are present when required; every spoken word, meaningful sound cue, speaker ID, and off-screen speech is represented as appropriate.
- Captions are timed to speech, not late; do not cover essential UI/product/legal information; remain readable over the background; stay within safe margins; and are not burned in if the platform/client requires sidecar captions.
- Subtitle translation preserves meaning, tone, names, numbers, legal claims, humor, and safety instructions. Do not accept a localization solely because it "sounds fluent."
- Audio description or text alternative exists when important visual-only information is necessary for comprehension.
- Alt text or extended descriptions exist for still images when the delivery context requires accessible images.
- Flashing/strobing sequences are tested or avoided. For public-facing video, treat fast flashes, saturated red flashes, and high-contrast repetitive patterns as a safety risk, not just a style choice.
Rights, provenance, and disclosure
Generated media QA must include origin and usage review. Do not present metadata as proof of truth; use it as one trust signal.
Documented facts:
- C2PA defines technical specifications for content provenance; a C2PA manifest can contain assertions, claims, signatures, and content bindings for media provenance.
- IPTC Digital Source Type includes
trainedAlgorithmicMediafor media created using generative AI andcompositeWithTrainedAlgorithmicMediafor composites containing generative-AI elements. - Google Merchant Center documentation says not to remove embedded metadata tags such as IPTC
DigitalSourceTypefrom AI-generated images; relevant NewsCodes includeTrainedAlgorithmicMedia. Verified 2026-07-10; platform requirements are volatile. - YouTube Help says creators must disclose AI-generated or meaningfully AI-altered content that appears realistic through the AI-use setting, and disclosed content can be labeled for viewers. Verified 2026-07-10; platform requirements are volatile.
- FTC advertising guidance does not create an "AI exemption" from truthful advertising rules; endorsements, testimonials, and product claims still need appropriate disclosure and substantiation.
Review:
- Source rights: stock license, music license, font license, logo permission, brand assets, user-provided media permission, and commercial usage terms.
- Likeness/voice: consent for synthetic or cloned voices, avatars, face swaps, public figures, employees, customers, and minors.
- Claims: product performance, before/after imagery, health/financial/legal benefits, prices, availability, endorsements, awards, comparative claims, and environmental claims.
- Provenance metadata: preserve C2PA/IPTC/XMP where required; record if metadata was stripped by editing or platform export; do not claim "verified real" just because a credential exists.
- Disclosures: add human-visible AI/synthetic disclosure when the platform, law, client policy, or audience deception risk requires it. For realistic people/events, political content, news-like scenes, ads, and endorsements, escalate rather than guessing.
- Model/provider metadata: capture provider, model, date, prompt lineage, reference sources, post tools, and revision history. If exact model version is unavailable, state that limitation.
Safety and policy review
Apply the relevant model, platform, legal, and client policies. Platform facts change; verify at delivery time for regulated or high-risk content.
Flag or reject:
- Deceptive realistic synthetic people, events, endorsements, evidence, or news scenes.
- Non-consensual intimate imagery, sexualized minors, harassment, hate, extremist persuasion, self-harm encouragement, or instructions for wrongdoing.
- Medical, financial, legal, political, housing, employment, or credit claims without substantiation and required disclaimers.
- Misrepresentation of a product's capabilities, certification, price, availability, environmental impact, or safety.
- Privacy leaks: faces, license plates, addresses, screens, documents, account details, patient/customer data, location traces, or hidden metadata.
- Localization harms: culturally offensive imagery, mistranslated warnings, wrong units, taboo gestures, or regionally illegal claims.
Revision triage
Use severity to decide whether to accept, revise, regenerate, or escalate.
Critical: must not ship. Examples: wrong product/claim, missing rights or consent, deceptive synthetic endorsement, disallowed safety/policy content, inaccessible required captions, unsafe flashing, corrupted final file, unintelligible core dialogue, wrong legal copy, or platform-required disclosure missing.
Major: revise before normal release unless stakeholder explicitly accepts risk. Examples: visible face/hand/logo defect, lip-sync drift, continuity break that harms comprehension, caption timing failures, audio clipping, noisy mix, bad localization, wrong brand color on key asset, or noticeable compositing seam.
Minor: fix if efficient; can ship with documented acceptance for low-stakes drafts. Examples: small background artifact, slight caption line break awkwardness, non-critical metadata typo, tiny compression artifact not visible at target size.
Observation: record without blocking. Examples: "AI-generated background texture visible on paus
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: calesthio
- Source: calesthio/generative-media-skills
- License: MIT
- Homepage: https://github.com/calesthio/OpenMontage
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.