# Talking Head Podcast Recut

> Provider-independent production workflow for recutting long talking-head, podcast, interview, webinar, panel, lecture, livestream, founder call, or customer-call footage into short clips, highlight reels, trailers, audiograms, captioned vertical videos, and social cutdowns. Use when an agent must ingest source recordings, transcribe and diarize speakers, select quotes without misrepresentation, p…

- **Type:** Skill
- **Install:** `agentstack add skill-calesthio-generative-media-skills-talking-head-podcast-recut`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [calesthio](https://agentstack.voostack.com/s/calesthio)
- **Installs:** 0
- **Category:** [Developer Tools](https://agentstack.voostack.com/c/developer-tools)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [calesthio](https://github.com/calesthio)
- **Source:** https://github.com/calesthio/generative-media-skills/tree/main/skills/production/content-formats/talking-head-podcast-recut
- **Website:** https://github.com/calesthio/OpenMontage

## Install

```sh
agentstack add skill-calesthio-generative-media-skills-talking-head-podcast-recut
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Talking-Head Podcast Recut

Use this skill to turn existing speech-led media into truthful, useful derivative edits. The central obligation is editorial integrity: every cut must remain faithful to what the speaker meant in the source recording, and every commercial, synthetic, legal, or rights risk must be surfaced before publishing.

This is provider-independent. Use whatever local or hosted tools are available for transcription, diarization, editing, captions, audio repair, motion graphics, and export, but keep the workflow, ledgers, and review standards below intact.

## Operating principle

Treat a recut as a quotation with pictures, not as raw material to rewrite.

- Do not create a clip whose apparent claim, emotional tone, chronology, or endorsement differs from the source.
- Do not use reaction shots, cutaways, captions, titles, or music to imply agreement, conflict, surprise, confidence, causality, or product endorsement that the original source does not support.
- Do not publish synthetic voice, face, avatar, or reconstructed dialogue without explicit approval from the rights holder and the depicted/simulated person.
- Do keep enough context for the clip to be understandable without requiring the audience to know the full episode.
- Do keep an audit trail from final frame back to source timecode.

## Intake: refuse to cut blind

Before editing, collect or infer a production brief. If any high-risk item is unknown, pause or mark it for approval.

Required brief fields:

- Source files or URLs, recording date, event name, speakers, host, guest, and known restrictions.
- Desired deliverables: single clip, clip batch, highlight reel, trailer, audiogram, captioned vertical video, LinkedIn cutdown, YouTube chapter clip, ad creative, internal sales enablement, or other use.
- Target platforms, aspect ratios, max duration, thumbnail/title/copy needs, caption style, language, brand guidelines, and review deadline.
- Intended use: organic editorial, paid ad, sales enablement, testimonial, investor update, course material, news, internal training, or archival.
- Rights status: who owns source footage, music, slides, screen shares, logos, transcript, and speaker likeness.
- Approval path: editor, producer, client, speaker/guest, legal/compliance, brand, accessibility, and final publisher.
- Risk notes: regulated topics, minors, medical/financial/legal claims, employment claims, politics/elections, customer testimonials, unreleased product information, confidential call participants, or sponsor/affiliate relationships.

Create a production ledger before cutting:

```text
source_id:
  path_or_url:
  owner:
  recording_date:
  speakers:
  permitted_uses:
  restrictions:
  transcript_path:
  diarization_path:
  generated_assets:
  approvals_required:
```

## Source ingest and transcript workflow

1. Preserve the original. Work from a duplicate or proxy. Record source file names, checksums if available, duration, frame rate, resolution, channels, and any embedded captions or metadata.
2. Extract high-quality audio for transcription. If the recording has separate mics/tracks, keep them separate for repair and diarization.
3. Generate a timestamped transcript with speaker diarization. If diarization is weak, use speaker names only after manual verification.
4. Create a speaker map: name, role, organization, pronunciation, lower-third wording, consent status, and allowed uses.
5. Mark unusable or restricted sections: private information, off-record sections, unapproved names, copyrighted clips, embargoed product material, sponsor reads that cannot be repurposed, audience Q&A with unconsented participants, or low-confidence transcript spans.
6. Normalize transcript text lightly for reading, but keep a verbatim layer for quote verification. Do not "clean up" meaning.
7. Build searchable markers for topic, emotion, claim, proof, story beat, objection, audience question, and repeated filler.

Useful transcript markers:

```text
[00:13:42.2 - 00:14:21.8] Speaker: Maya Chen
Topic: onboarding friction
Claim: "activation doubled after we removed the first five setup steps"
Context before: discussing enterprise admins, not all users
Context after: says result was a pilot, not general launch
Risk: metric requires client approval
Clip potential: strong, but title must say "pilot" or avoid universal claim
```

## Rights, consent, and disclosure gate

This skill does not provide legal advice. It provides escalation triggers for the agent. When in doubt, ask the client, rights owner, counsel, or participant before cutting or publishing.

Escalate before production or publication when:

- The source recording was private, semi-private, internal, customer-facing, webinar-gated, or captured in a jurisdiction or setting where consent expectations are unclear.
- A participant did not explicitly agree to public repurposing, paid promotion, advertising, testimonial use, voice cloning, face replacement, avatar generation, translation dubbing, or synthetic extension.
- The clip uses a customer, employee, expert, creator, or influencer in a way that could be an endorsement or testimonial.
- The output will be used as an ad, sales asset, fundraising asset, political communication, medical/financial/legal advice, employment/recruiting claim, or investor claim.
- The recut includes third-party music, TV/movie/game footage, slides, screenshots, artwork, conference footage, audience Q&A, user-generated content, or logos beyond the source owner's rights.
- The source is newsworthy or critical commentary and the edit relies on fair use rather than permission.
- The title, thumbnail, caption, or intro card makes a stronger claim than the speaker made.
- The speaker discusses another person, customer, employer, patient, student, or confidential situation.

Commercial and endorsement disclosure:

- Documented fact, verified 2026-07-11: FTC endorsement guidance treats material connections as needing clear disclosure when an endorsement can be attributed to an advertiser or marketer. Use clear, hard-to-miss disclosure for sponsorships, paid relationships, free products, affiliate relationships, or brand control. Source: FTC endorsement guidance and Disclosures 101.
- Documented platform examples, verified 2026-07-11: YouTube has paid-promotion declarations; TikTok requires commercial-content disclosure for content promoting a brand/product/service; Instagram branded content policies require use of branded-content tools for branded content; LinkedIn says value-exchange posts must be labeled as brand partnerships. These platform facts are volatile; re-check platform help pages at production time.

Synthetic and altered-media disclosure:

- Documented platform examples, verified 2026-07-11: YouTube requires creators to disclose realistic altered or synthetic content in defined cases, and TikTok requires labeling AI-generated content that contains realistic images, audio, or video. Re-check the current upload workflow and policy text before publishing.
- If a synthetic narrator, cloned speaker voice, avatar, lip-sync, face swap, translated dub, AI filler word reconstruction, or generated b-roll depicts a real person or realistic event, require explicit approval and a disclosure plan.

Copyright and fair use:

- Documented fact, verified 2026-07-11: U.S. Copyright Office fair-use guidance says there is no fixed number of words, notes, or percentage that is automatically safe; fair use depends on circumstances. Do not present clip length as permission. Escalate permission/fair-use decisions to counsel or the rights owner.

Provenance:

- When available, preserve or attach provenance metadata. C2PA Content Credentials are a standards-based way to describe media origin and edit history, but metadata can be stripped by platforms or workflows; keep an internal ledger even when credentials are present.

## Quote selection: the context test

For every candidate quote, answer these before it can become a clip:

1. What is the speaker's exact claim in plain language?
2. What did the speaker say immediately before and after?
3. Would the speaker reasonably agree this clip represents their point?
4. Does the clip omit a qualification such as "in our pilot," "for this customer," "we think," "we failed first," or "not always"?
5. Does the hook/title/caption introduce a claim not present in the source?
6. Does the clip use music, reaction shots, zooms, or graphics to intensify emotion beyond the source?
7. Does this quote need factual, legal, sponsor, customer, or speaker approval?

Prefer quotes that:

- Contain one clear idea, tension, lesson, story turn, or useful answer.
- Start quickly or can be introduced with a truthful setup card.
- Have a natural ending, punchline, decision, proof, or invitation.
- Work without relying on the full episode.
- Include enough specifics to feel credible, but not so much private or regulated detail that approval becomes fragile.

Reject or repair quotes that:

- Depend on pronouns or missing referents that require misleading titles.
- Sound stronger when isolated than in context.
- Are transcript artifacts, jokes without context, sarcasm, or cross-talk.
- Include private data, medical/financial/legal advice, unpublished performance data, or third-party allegations without approval.
- Require sentence splicing that makes the speaker say a sentence they did not say.

## Clip structures

Use these as production heuristics, not mandatory formulas.

### Single insight vertical clip

- 0-2s: truthful hook from the speaker or a setup title.
- 2-8s: context needed to understand the quote.
- 8-35s: core insight, story, or answer.
- 35-50s: implication, lesson, or memorable line.
- Last 1-3s: clean end beat, branded outro, or platform-native CTA if approved.

### Two-speaker exchange

- Start with the question if it is short and necessary.
- Preserve the answer's target. Do not cut a question about one topic into an answer about another.
- Use speaker labels and lower thirds, especially in vertical crop where faces may alternate.
- Avoid reaction shots from a different moment unless labeled or neutral enough not to imply a response.

### Highlight reel or trailer

- Organize around a promise: "three hard lessons from scaling support," not "random best moments."
- Use a spine: problem -> tension -> proof -> consequence -> invite.
- Maintain chronological clarity if chronology matters. If using non-linear montage, do not imply events happened in the order shown.
- Use on-screen topic cards to prevent decontextualized quote stacking.

### Audiogram

- Use when the visual source is weak, audio-only, remote-call quality, or a podcast platform needs shareable video.
- Include waveform only if it helps motion; do not let it compete with captions.
- Use host/guest photos, title, episode art, and topic cards with rights-confirmed assets.
- Ensure captions and speaker IDs carry the experience for muted viewing.

### Customer/founder/testimonial clip

- Require release/approval for testimonial use.
- Preserve the actual experience and scope. Do not generalize a pilot/customer quote into a universal result.
- Add disclosure if there is a material connection or incentive.
- Submit claims, logos, customer name, metrics, and title/thumbnails for client/legal approval.

## Edit craft

### A-roll assembly

- Cut for clarity before pace. Remove dead air, false starts, and repetitions only when doing so preserves meaning.
- Keep breaths and micro-pauses where they support authenticity or comprehension.
- Use jump cuts confidently for social clips, but hide distracting cuts with scale changes, B-roll, caption beats, or J/L cuts when the face movement is jarring.
- Never combine words from separate contexts to create a new sentence unless it is purely grammatical cleanup and the result is transcript-verifiable.
- Use room tone or ambience under repaired gaps. Avoid hard digital silence.
- If cutting between speakers, preserve conversational rhythm. Do not make an answer appear more immediate, hostile, or enthusiastic than it was.

### B-roll, screen, slides, and graphics

- B-roll should clarify, evidence, or pace the spoken point. Do not use generic filler that changes the topic.
- Use product footage, slides, charts, screenshots, or customer logos only when rights and freshness are confirmed.
- Use text callouts for names, terms, numbers, and transitions; avoid restating every spoken word outside captions.
- Lower thirds should include speaker name and role as approved; avoid inflated titles.
- If adding stock or generated visuals, record provenance and ensure they do not depict real people/events as if they were source footage.

### Cropping and reframing

- Reframe from the highest-resolution source available.
- For two-person remote calls, consider dynamic crop switching only when it improves comprehension and does not feel like fake reaction editing.
- Keep faces, hand gestures, slide text, and captions inside platform safe areas.
- Avoid excessive punch-ins on sensitive, emotional, or serious statements; they can editorialize the tone.

## Captions and accessibility

Documented facts, verified 2026-07-11:

- WCAG 2.2 Success Criterion 1.2.2 requires captions for prerecorded audio content in synchronized media, with exceptions for media alternatives.
- W3C understanding guidance says captions include dialogue, identify who is speaking, and include non-speech sound information needed to understand the content.
- WCAG includes transcript and audio-description/media-alternative requirements for some media contexts.

Caption workflow:

1. Generate captions from the verified transcript, not from a rough draft.
2. Correct names, jargon, product terms, numbers, and homophones manually.
3. Include speaker IDs when the speaker is not visually obvious or when the clip is audio-only/audiogram style.
4. Include meaningful non-speech audio when relevant, such as "[laughter]," "[applause]," or "[door closes]" if it affects meaning.
5. Keep caption lines readable. Use natural phrase breaks; do not split names, numbers, or dependent clauses awkwardly.
6. Use high contrast, sufficient size, and safe placement. Avoid hiding captions behind platform UI, lower thirds, progress bars, or auto-generated overlays.
7. Deliver both burned-in captions for social clips and sidecar captions (SRT/VTT) when the platform or client can use them.
8. For multilingual captions or translated cutdowns, verify translation against the source meaning and get approval for sensitive claims.

## Audio cleanup and loudness

Documented facts, verified 2026-07-11:

- ITU-R BS.1770 specifies algorithms for measuring programme loudness and true-peak level.
- EBU loudness guidance describes R 128 normalization around -23 LUFS for broadcast-style workflows.
- ATSC A/85 provides loudness guidance for television and streaming-media service providers.

Production heuristics for social and web cutdowns:

- Prioritize intelligibility of speech over loud music or heavy compression.
- Repair in this order: remove hum/clicks/plosives where possible, reduce broadband noise conservatively, de-reverb if needed, balance speakers, EQ for clarity, compress lightly, limit true peaks, then loudness-normalize.
- For web/social delivery where no client spec exists, a practical target is often around -14 to -16 LUFS integrated with true peak at or below -1 dBTP. Treat this as a platform-era heuristic, not a legal or universal standard.
- For broadcast, OTT, paid media, or client delivery, use the required spec exactly; do not substitute social loudness targets.
- Check the final export, not only the timeline mix. Captions, stingers, end cards, and music beds can change measured loudness.

Audio QA:

- No clipped syllables at edit points.
- No denoise warble or underwater artifacts on speech.
- Speakers balanced within the clip.
- Music ducks under dialogue and never hides required disclosure.
- Endings do no

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [calesthio](https://github.com/calesthio)
- **Source:** [calesthio/generative-media-skills](https://github.com/calesthio/generative-media-skills)
- **License:** MIT
- **Homepage:** https://github.com/calesthio/OpenMontage

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-calesthio-generative-media-skills-talking-head-podcast-recut
- Seller: https://agentstack.voostack.com/s/calesthio
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
