Install
$ agentstack add skill-tofuswang-video-subtitle-skill-video-subtitle ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
/video-subtitle — Video Subtitle Tool
Transcribe video audio to SRT subtitles using a local Whisper model, then burn them into the video.
User-invocable
When the user types /video-subtitle, run this skill.
Prerequisites
This skill requires two tools. Check availability before proceeding:
1. mlx-whisper (Apple Silicon) or openai-whisper
Detect which is available:
which mlx_whisper whisper 2>/dev/null
If neither is found, guide the user to install one:
- Apple Silicon Mac (recommended):
pip install mlx-whisper - Any platform:
pip install openai-whisper
2. ffmpeg with libass
Check for libass support:
ffmpeg -filters 2>&1 | grep subtitles
If no output (libass missing), the subtitles filter is unavailable:
- macOS:
brew install ffmpeg-full(keg-only, use/opt/homebrew/opt/ffmpeg-full/bin/ffmpeg) - Linux:
sudo apt install ffmpeg(usually includes libass) - Fallback: If the user has regular ffmpeg without libass, soft-embed the SRT as a subtitle track instead of burning in.
Whisper model selection
Discover cached models:
ls ~/.cache/huggingface/hub/ 2>/dev/null | grep whisper
Prefer the largest available model. If none cached, use:
- mlx-whisper:
mlx-community/whisper-large-v3-mlx - openai-whisper:
large-v3
Modes
/video-subtitle path/to/video.mov→ Full pipeline (transcribe + correct + burn)/video-subtitle srt path/to/video.mov→ SRT only (transcribe, no burn)/video-subtitle burn path/to/video.mov path/to/sub.srt→ Burn only (use existing SRT)/video-subtitle(no args) → Interactive, ask for video path
Full Pipeline
Step 1: Verify video
If no path provided, ask with AskUserQuestion. Confirm the file exists and read metadata:
ffmpeg -i "" 2>&1 | grep -E "Duration|Stream.*(Video|Audio)"
Report duration, resolution, and file size to the user.
Step 2: Extract audio
ffmpeg -i "" -vn -acodec pcm_s16le -ar 16000 -ac 1 "/_audio.wav" -y
16kHz mono WAV is the optimal input format for Whisper.
Step 3: Confirm language
Ask the user with AskUserQuestion. Common options:
| Language | Flag | |----------|------| | Chinese | --language zh | | English | --language en | | Japanese | --language ja | | Auto-detect | omit --language |
Step 4: Transcribe
mlx-whisper:
mlx_whisper --model "" --language --output-format srt --output-dir "" ""
openai-whisper:
whisper "" --model large-v3 --language --output_format srt --output_dir ""
Tell the user: "Transcribing, please wait..."
Step 5: Clean up SRT
Read the generated SRT and fix:
- Remove hallucinations: Delete entries with zero-duration timestamps (e.g.
00:00:57,640 --> 00:00:57,640) or empty text - Fix proper nouns: Whisper commonly misrecognizes brand names and technical terms — use conversation context to correct
- Merge fragments: Combine entries shorter than 0.5s that are semantically connected
- Renumber: Run this Python snippet to renumber all entries:
python3 -c "
import re
with open('', 'r') as f:
content = f.read()
blocks = re.split(r'\n\n+', content.strip())
out = []
for i, block in enumerate(blocks, 1):
lines = block.strip().split('\n')
if len(lines) >= 2:
lines[0] = str(i)
out.append('\n'.join(lines))
with open('', 'w') as f:
f.write('\n\n'.join(out) + '\n')
"
Present the corrected SRT to the user and ask if further edits are needed before burning.
Step 6: Burn subtitles
Determine the correct ffmpeg binary (prefer ffmpeg-full if regular ffmpeg lacks libass):
FFMPEG="ffmpeg"
if ! ffmpeg -filters 2>&1 | grep -q subtitles; then
if [ -x /opt/homebrew/opt/ffmpeg-full/bin/ffmpeg ]; then
FFMPEG="/opt/homebrew/opt/ffmpeg-full/bin/ffmpeg"
fi
fi
Then burn:
$FFMPEG \
-i "" \
-vf "subtitles=:force_style='FontSize=24,PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,Outline=2,MarginV=30'" \
-c:v libx264 -preset medium -crf 23 \
-c:a aac -b:a 128k \
-pix_fmt yuv420p \
"/_subtitled.mp4" -y
force_style options:
| Param | Default | Description | |-------|---------|-------------| | FontSize | 24 | Font size | | PrimaryColour | &H00FFFFFF | White text | | OutlineColour | &H00000000 | Black outline | | Outline | 2 | Outline thickness | | MarginV | 30 | Bottom margin |
Important: The SRT file path must not contain special characters that confuse the ffmpeg filter parser. Use relative paths or symlink if needed.
Step 7: Done
Report the output file path and size. Suggest open to preview. Remind the user that the intermediate _audio.wav file can be deleted.
Troubleshooting
| Problem | Cause | Fix | |---------|-------|-----| | No such filter: 'subtitles' | ffmpeg missing libass | Use ffmpeg-full or install libass | | No such filter: 'drawtext' | ffmpeg missing libfreetype | Use ffmpeg-full | | Duplicate/empty SRT entries | Whisper hallucination | Clean in Step 5 | | Chinese chars show as boxes | Missing CJK font | macOS: auto-fallback to PingFang. Linux: apt install fonts-noto-cjk | | Error parsing filterchain | Special chars in SRT path | Use relative path, avoid spaces |
Quality Rules
- Faithful transcription: Only transcribe what was actually said — do not add or rewrite meaning
- Proper noun correction: Fix common Whisper misrecognitions, but confirm with the user
- Subtitle rhythm: Max two lines per entry, max ~20 CJK characters per line
- Clean up: Remind user to delete intermediate files (_audio.wav) after completion
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: Tofuswang
- Source: Tofuswang/video-subtitle-skill
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.