Install
$ agentstack add skill-vibetool-talking-head-video-talking-head-video ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Talking-Head Explainer Video
Build a finished 口播 video (vertical 9:16 by default, landscape 16:9 fully supported — see 画幅 section) where the speaker is full-frame by default, and during each explanation beat the speaker shrinks into a circular avatar in the top-left while an animated knowledge graphic (chart / data cards / comparison / flow / big-number) fills the screen. Subtitles are burned in and corrected against the user's verbatim script.
This skill captures a real working pipeline. Read this whole file first, then the references as needed. The design conventions exist for concrete reasons — keep them unless the user asks otherwise.
The core idea (why it's built this way)
The original talking-head video is the spine — its audio plays untouched the whole time, so sync is never an issue. You only overlay graphics + a shrunk avatar during "knowledge windows"; everywhere else it's just the full face. The graphics render as ProRes 4444 (alpha) with a short fade at their edges, so they crossfade onto the face instead of cutting to black.
full face ──┐ ┌── graphic full-frame + avatar top-left ──┐ ┌── full face
│ cross │ (one "knowledge window") │ cross │
subtitles burned over EVERYTHING, the whole time
Prerequisites (check / install once)
- ffmpeg with libass + prores_ks (Homebrew ffmpeg has these)
- faster-whisper:
python3 -m pip install faster-whisper - Node + Remotion: the template in
assets/remotion/—npm installinside it.
Remotion renders via a headless Chrome it manages.
- A CJK font libass can find (macOS: PingFang SC works out of the box).
Workflow
Set up a working dir (e.g. ~/Downloads/_build/). Copy assets/remotion/ into it as remotion/. Then:
1. Inspect the source
ffprobe the clip for resolution/duration/fps. Vertical 1080×1920 @ 30fps and landscape 1920×1080 @ 30fps are both first-class (see 画幅 section for the parameter set of each). 开工前先和用户确认目标画幅(通常跟源视频走, 但要问一句,别默默定;用户已明说的不再问). Extract a frame (ffmpeg -ss 2 -i src.mp4 -frames:v 1 f.jpg) and look at where the face sits — you need this for the avatar crop.
开工前先问用户(硬规则)
制作视频前,必须和用户确认两个选项,给推荐但由用户拍板;用户已明说的不重复问:
- 画幅:横屏 16:9 还是竖屏 9:16(参数表见下节);
- 图卡风格:darkcard / watercolor / inkline(见「图卡风格」节)。
图卡风格 · Graphics styles(三选一)
| 风格 | 一句话 | 适合 | 字体(fonts.ts 现成栈) | |---|---|---|---| | darkcard 深色科技卡(默认) | 深底、辉光、边框卡片、大数字 | 价格/参数/对比/新闻/评测 | 得意黑(DARKTITLE)+ DIN(DARKNUM)+ PingFang | | watercolor 纸上水彩 | 米黄纸、水彩物件、铅字标题、物理点睛 | 生活类比/故事感/章节卡 | 汇文明朝体(WCCN)+ Special Elite(WCEN) | | inkline 白纸简笔 | 纯白、细线"画出来"、大字+灰注、诗意收尾 | 原理/定义/一句话洞察 | 霞鹜文楷双字重(INK) |
工艺规范(色板/滤镜/动效语法/选择速查)在 references/styles.md,写组件前先读。darkcard 的六个 archetype 在 references/design-system.md + assets/remotion/src/;watercolor/inkline 的基础件在 src/watercolor.tsx / src/inkline.tsx;三风格各有完整示例在 src/examples/ (同一概念「边际成本」三种表达,当模板抄结构)。字体接线:bash scripts/setup_fonts.sh (自带 6 个免费字体;Anthropic 官方字体从本机 Claude.app 自动检测,禁止随仓库/产物分发)。 通用红线:文字永不加做旧/毛边滤镜,滤镜只作用于图形线条。
画幅 · Orientation parameter sets (both proven on real videos)
| 参数 | 竖屏 9:16 (1080×1920) | 横屏 16:9 (1920×1080) | |---|---|---| | Remotion Composition | width 1080, height 1920 | width 1920, height 1080 | | 头像直径 / overlay 位置 | 300px / 40:120 | 380px / 44:44 | | make_avatar.sh 参数 | CROP≈1040:1040:0:250 SIZE=300 | CROP≈780:780:580:20 SIZE=380(居中脸源) | | 字幕 ASS PlayRes / MarginV | 1080×1920 / 210 | 1920×1080 / 90(generate_subs.py 传 --land) | | 图形安全区 | 避开左上头像 + y≳1540 字幕带 | 避开左上头像(x out.png --frame=`
Critical layout rule: keep all graphic content above y≈1540. The bottom ~360px is reserved for subtitles. Put payoff callouts at bottom: 380–430, not bottom: 150. (This was the #1 bug on the first build.)
5. Render graphics to alpha clips
bash scripts/render_graphics.sh remotion clips ...
Renders ProRes 4444 .mov (alpha) into clips/. Alpha is required for the edge crossfade.
6. Make the circular avatar
bash scripts/make_avatar.sh source.mp4 avatar.mov 1040:1040:0:250 0x22D3A8
Tune the crop so the whole head incl. chin sits in the circle — wider crop = smaller face. Validate on one frame first (the script prints where it wrote). Common mistake: too tight a crop clips the chin.
7. Subtitles (corrected)
python3 scripts/generate_subs.py asr.json subs.ass
Edit CORRECTIONS in the script per video — whisper mis-hears brand names / English terms / 同音字 (e.g. 施实验→湿实验, 礼来, Indeed, Verve, 胰腺癌, PPT). Skim asr.json and add the wrong→right pairs before running.
8. Composite → final
Write segments.json (source, avatar, subtitles, output, and the windows), then:
python3 scripts/build_composite.py segments.json
It generates the ffmpeg filtergraph and renders final.mp4 (graphics overlaid in their windows + avatar top-left during windows + subtitles burned + original audio). Use --dry-run to inspect the filter first.
9. QA (always)
Extract a frame from the final mp4 at the middle of every segment type and look at each:
ffmpeg -ss -i final.mp4 -frames:v 1 qa_.jpg
Check, per frame: avatar shows the whole face and sits in the clear top-left zone; subtitle doesn't collide with graphic content; graphic content isn't cut off; no black flash at window boundaries (mid-crossfade dark is fine). Fix, re-render only the affected clip(s), re-composite.
Editorial conventions worth keeping
- Don't make it all graphics. The face carries the hook and the emotional
conclusion. Graphics are for facts/data/structure.
- One subtitle correction pass is mandatory — uncorrected ASR (wrong brand
names, 同音字) instantly reads as sloppy.
- Watch the narrative framing in graphics. If the speaker's claim is a slight
oversimplification you can't change (it's recorded audio), make the graphic text accurate so it doesn't contradict knowledgeable viewers. Likewise, if a segment risks an unintended read (e.g. "AI replaces humans"), add a bridging line in the graphic that steers toward the speaker's actual point.
Reusable assets in this skill
assets/remotion/— full Remotion project:theme.tsx(design system) + 6
archetype components + Root.tsx. Copy into the project and adapt.
scripts/transcribe.py— faster-whisper word timingsscripts/generate_subs.py— ASR → corrected styled.assscripts/make_avatar.sh— circular alpha PiP avatar (parameterized crop/ring)scripts/render_graphics.sh— batch ProRes-4444 renderscripts/build_composite.py— segments.json → ffmpeg master compositescripts/segments.example.json— the segment-map schemareferences/design-system.md— palette, tokens, avatar/subtitle geometry, and
the graphic archetype catalog. Read before designing graphics.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: Vibetool
- Source: Vibetool/talking-head-video
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.