Install
$ agentstack add skill-erictechpro-startup-claude-skills-scoop Open-source listing — not yet scanned by AgentStack. Follow the source repository for install instructions.
Security review
⚠ Flagged1 finding(s); flagged for manual review. · v0.1.0 How review works →
- • Prompt-injection patterns
- • Secret / credential exfiltration
- • Dangerous shell & filesystem operations
- • Untrusted network calls
- • Known-malicious package signatures
- high Dangerous shell/eval execution.
What it can access
- ✓ Network access No
- ● Filesystem access Used
- ● Shell / process execution Used
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Scoop — Content Gap Analysis
Extracts ALL comments from competitor YouTube videos (and optionally your own), identifies content gaps ranked by audience engagement, and synthesizes a content strategy.
Why this skill exists
Content-gap analysis done by hand is mostly dull categorization — reading hundreds of comments, clustering them, counting likes. It's the kind of work where the model is great at the qualitative part (grouping a comment under the right theme, choosing keywords for a gap) and unreliable at the quantitative part (adding up likes, merging counts across videos). An earlier version of this skill let the model do both, and the arithmetic came out wrong in ways the user couldn't audit. This version splits the work: the model classifies, Python scores. Every score in the final output is traceable back to the comment IDs that produced it.
Setup (one-time)
The underlying toolkit eagerly imports apify_client from its package __init__ even though this skill doesn't use Apify. If you haven't installed it:
pip3 install apify-client
Do this once per environment. It's a known toolkit quirk.
Transcription requires mlx_whisper (Apple Silicon):
pipx install mlx-whisper
The model (mlx-community/whisper-medium-mlx, ~1.5GB) downloads automatically on first run.
Pipeline Overview
┌─────────┐ ┌─────────┐ ┌──────────┐ ┌──────────┐ ┌─────────┐ ┌──────────┐
│ 1. │──>│ 2. │──>│ 3. │──>│ 3.5 │──>│ 4. │──>│ 5. │
│ Extract │ │ Classify│ │ Verify │ │ Channel │ │ Score │ │ Narrate │
│ videos │ │ gaps │ │ vs trans-│ │ gap │ │ gaps │ │ strategy │
│ (JSON) │ │ (JSON) │ │ cript │ │ check │ │ (md) │ │ (prose) │
└─────────┘ └─────────┘ └──────────┘ └──────────┘ └─────────┘ └──────────┘
Python Model Python Python Python Model
side side side side side side
Steps 1, 3, 3.5, 4 are deterministic Python. Steps 2 and 5 are the model's creative work. The separation is the whole point — it means scores are reproducible and the strategy narrative is grounded in facts Python verified.
All intermediate files live in /tmp/:
/tmp/scoop_videos.json— extracted data for all videos/tmp/scoop_gaps.json— classified gaps (written by model, cross-check appendscovered_in_transcript)/tmp/scoop_scored.md— rendered ranked tables (fed back into the narrative)
Final narrative output lives in output/scoop/.md — the user-facing deliverable produced in Step 5. output/ is gitignored at the repo root, so these files stay out of version control by design.
Step 1 — Extract Videos
1a. Collect URLs
Ask the user two things via AskUserQuestion:
> "What competitor video URLs do you want to analyze? (paste one or more YouTube URLs)"
Then:
> "Do you have your own video on this same topic? If yes, paste the URL. If not, type N/A."
1b. Extract and write a single combined JSON
Run ONE Python block that extracts every URL and writes one combined file. Role-tag each video (competitor1, competitor2, …, or own) so later steps can distinguish them.
import json, sys, subprocess, os
from yt_toolkit.extractors.video import VideoExtractor
URLS = [
("competitor1", "https://www.youtube.com/watch?v=..."),
("competitor2", "https://youtu.be/..."),
("own", "https://youtu.be/..."), # omit tuple if user said N/A
]
extractor = VideoExtractor()
videos = []
for role, url in URLS:
try:
data = extractor.extract(url, all_comments=True, save_to_file=False)
except Exception as e:
print(f"[skip] {role} {url}: {e}", file=sys.stderr)
continue
# --- Local transcription via mlx_whisper ---
vid = data.get("video_id", role)
audio_path = f"/tmp/scoop_{vid}.mp3"
transcript = {"available": False, "text": "", "segments": [], "language": ""}
try:
# Download audio
subprocess.run(
["yt-dlp", "-x", "--audio-format", "mp3", "--audio-quality", "0",
"-o", audio_path, "--no-playlist", url],
check=True, capture_output=True, text=True,
)
# Transcribe locally (auto-detect language)
subprocess.run(
["mlx_whisper", audio_path,
"--model", "mlx-community/whisper-medium-mlx",
"--output-format", "json", "--output-dir", "/tmp"],
check=True, capture_output=True, text=True,
)
json_path = audio_path.rsplit(".", 1)[0] + ".json"
if os.path.exists(json_path):
with open(json_path) as jf:
w = json.load(jf)
transcript = {
"available": True,
"text": w.get("text", ""),
"segments": [
{"text": s["text"], "start": s["start"],
"duration": round(s["end"] - s["start"], 3)}
for s in w.get("segments", [])
],
"language": w.get("language", "en"),
}
os.remove(json_path)
else:
print(f"[warn] {role}: whisper JSON not found at {json_path}", file=sys.stderr)
except subprocess.CalledProcessError as e:
print(f"[warn] {role}: transcription failed: {e}", file=sys.stderr)
except Exception as e:
print(f"[warn] {role}: transcription error: {e}", file=sys.stderr)
finally:
if os.path.exists(audio_path):
os.remove(audio_path)
data["transcript"] = transcript
# --- End local transcription ---
data["role"] = role
videos.append(data)
# Truncation guard — YouTube API pagination can silently stop short.
got = len(data.get("comments", []))
expected = data.get("statistics", {}).get("comments", 0)
if expected and abs(got - expected) / expected > 0.05:
print(f"[warn] {role}: extracted {got} comments, metadata says {expected} — possible truncation", file=sys.stderr)
stats = data.get("statistics", {})
tr = data.get("transcript", {})
print(f"{role}: {data['title'][:60]} | views {stats.get('views','?'):,} | "
f"comments got={got} meta={expected} | transcript={'yes' if tr.get('available') else 'no'}")
with open("/tmp/scoop_videos.json", "w") as f:
json.dump(videos, f)
print(f"Saved {len(videos)} videos to /tmp/scoop_videos.json")
Important: always pass all_comments=True. The default 20 is never enough — the whole value of Scoop is reading the long tail.
If extraction fails for one URL, continue with the rest. Don't abort. If transcription fails for one video, the video is still included (with transcript.available = False) — comments are the primary data, transcripts are for cross-checking.
1c. Present the extraction summary
After running the block, present a table to the user:
| # | Video Title | Channel | Views | Comments | Transcript | Role |
|---|------------------------------------------|----------------|---------|----------|------------|-------------|
| 1 | GSD vs Superpowers vs Claude Code | Chase AI | 26,453 | 96 | yes | competitor1 |
| 2 | This One Plugin Just 10x'd Claude Code | Nate Herk | 60,029 | 102 | yes | competitor2 |
| 3 | Claude Code + SUPERPOWERS = ... | Eric Tech | 27,409 | 10 | yes | own |
If a truncation warning fired, surface it here. If any transcript is missing, flag it — the cross-check in Step 3 will skip that video.
Step 2 — Classify Comments into Gaps
This is where the model earns its keep. Read /tmp/scoop_videos.json (or the summary you already produced in context) and cluster comments into content gaps — themes the audience keeps bringing up that represent unaddressed (or poorly addressed) demand.
What counts as a gap
- Content gap — a topic the commenter wants covered.
- Frustration — criticism of what the video did/didn't do (needs a different tone of response video, so we tag it).
- Follow-up request — explicit "please make a video on X."
- Alternative tool mentioned — a tool the audience keeps naming that the video ignored.
- Methodology criticism — what the audience thinks was done wrong.
These categories are lenses, not output buckets. Every gap gets represented in gaps.json with the same schema below.
Rules for classification
- Global, not per-video. Classify across all videos in a single pass. If the same theme shows up in V1 and V2, it's ONE gap with two
video_appearancesentries — not two separate gaps. - Cite comment IDs. For each gap, list every comment you're using as evidence. This is how scoring stays honest. Split
comment_ids(top-level) fromreply_comment_ids(nested replies) — the scoring script weights them differently. - Exclude creator self-comments. The YouTube API returns author display names (e.g.
@Chase-H-AI) which don't exact-match the channel display title (e.g.Chase AI), so naive string matching fails. Instead, identify the creator's handle per video by looking for pinned/promo comments ("Full courses + unlimited support: ...", "Get the Claude Code Masterclass: ..."), and record them ingaps.jsonascreator_handles_by_video: {video_id: ["@handle"]}. The scoring script reads this map and drops any cited comment whose author matches. Creators replying to their own comments aren't audience demand. - Keywords are search terms, not labels. Pick 3–6 distinct phrases that a transcript would contain if the video addressed this gap. Include synonyms. Bad:
["quality"]. Good:["code quality", "maintainability", "architecture", "test coverage", "production-ready"]. Step 3 uses these for a literal transcript search — vague keywords produce vague results. - Frustration flag. Set
frustration: truewhen the gap is driven by criticism of the existing video ("you missed X," "this is wrong"). Setfalsewhen it's demand or curiosity ("please make a video on X," "how do you do Y?"). One boolean, nothing more nuanced — boundary cases go tofalse. - One comment can cite multiple gaps. That's fine. The scoring script will log it so you can spot over-categorization.
Write the JSON
Write /tmp/scoop_gaps.json with the following schema:
{
"creator_handles_by_video": {
"celLbDMGy8w": ["@Chase-H-AI"],
"4XqVR6xI6Kw": ["@nateherk"]
},
"gaps": [
{
"gap_id": "g1",
"name": "Complex brownfield project comparison",
"description": "Audience wants testing on existing production codebases, not trivial greenfields.",
"frustration": true,
"keywords": ["brownfield", "complex project", "existing codebase", "production", "legacy"],
"video_appearances": [
{
"video_id": "celLbDMGy8w",
"comment_ids": ["UgxAbc123...", "UgyDef456..."],
"reply_comment_ids": ["UgzGhi789..."]
}
]
}
]
}
Aim for 6–12 gaps. Fewer than 6 usually means you're under-categorizing; more than 12 means you're splitting themes that belong together.
Step 3 — Transcript Cross-Check
This step catches the single biggest accuracy failure in gap analysis: ranking a "gap" as top missing content when the competitor actually covered it. A comment complaining "you never talked about code quality" is only useful evidence of a gap if the transcript confirms it. If the video dedicated 90 seconds to code quality and commenters still asked for more, that's a different kind of opportunity — "covered but shallow" — and it needs different framing in the final narrative.
This pass is a keyword search, not a depth judgment. It produces a boolean per gap per video.
Run this inline:
import json, re
with open("/tmp/scoop_videos.json") as f:
videos = json.load(f)
with open("/tmp/scoop_gaps.json") as f:
data = json.load(f)
by_id = {v["video_id"]: v for v in videos}
for gap in data["gaps"]:
coverage = []
for app in gap["video_appearances"]:
vid = app["video_id"]
v = by_id.get(vid)
if not v:
continue
tr = v.get("transcript") or {}
if not tr.get("available"):
coverage.append({"video_id": vid, "status": "NO_TRANSCRIPT", "hits": []})
continue
text = (tr.get("text") or "").lower()
segments = tr.get("segments") or []
hits = []
for kw in gap["keywords"]:
needle = kw.lower().strip()
if not needle:
continue
if needle in text:
# find approximate timestamp of first hit
for seg in segments:
if needle in (seg.get("text") or "").lower():
hits.append({"keyword": kw, "start": seg.get("start", 0)})
break
else:
hits.append({"keyword": kw, "start": None})
status = "COVERED" if hits else "NOT_COVERED"
coverage.append({"video_id": vid, "status": status, "hits": hits})
gap["coverage"] = coverage
with open("/tmp/scoop_gaps.json", "w") as f:
json.dump(data, f, indent=2)
# Human-readable summary
for gap in data["gaps"]:
statuses = [c["status"] for c in gap["coverage"]]
print(f"{gap['gap_id']} {gap['name'][:50]:50s} {statuses}")
How to read the output:
NOT_COVEREDin every video → genuine unaddressed gap. Top candidate for a new video.COVEREDin at least one video → the topic was touched on. Audience is asking for more depth, not introduction. Frame the recommendation accordingly ("deeper than X's 90-second mention").NO_TRANSCRIPT→ can't verify. Flag in the final output rather than silently treating it as NOT_COVERED.
Do NOT drop COVERED gaps from the output. Commenters asking for more depth on a covered topic is still real demand; it just needs different framing.
Step 3.5 — Channel Gap Check
Step 3 checks whether the analyzed video covered a gap. But the competitor might have a whole separate video about it on their channel. Without this check, Scoop cheerfully recommends "make a video about X" when the competitor already has one — just not the one you analyzed.
This step fetches the full video catalog for each competitor channel and keyword-searches titles and descriptions against the gap keywords.
Run this inline:
import json
from yt_toolkit.utils.youtube_api import YouTubeAPI
with open("/tmp/scoop_videos.json") as f:
videos = json.load(f)
with open("/tmp/scoop_gaps.json") as f:
data = json.load(f)
api = YouTubeAPI()
# Collect unique competitor channels (skip "own")
channels = {}
for v in videos:
if v.get("role", "").startswith("competitor") and v.get("channel_id"):
cid = v["channel_id"]
if cid not in channels:
channels[cid] = {
"channel_title": v.get("channel_title", ""),
"videos": api.get_channel_videos(cid, max_results=None),
}
print(f"Fetched {len(channels[cid]['videos'])} videos from {channels[cid]['channel_title']}")
# For each gap, check if any channel video title/description matches the keywords
for gap in data["gaps"]:
channel_coverage = []
for cid, ch in channels.items():
matching = []
for cv in ch["videos"]:
title = (cv.get("title") or "").lower()
desc = (cv.get("description") or "").lower()
searchable = title + " " + desc
for kw in gap["keywords"]:
if kw.lower().strip() in searchable:
matching.append({"video_id": cv["video_id"], "title": cv.get("title", "")})
break # one keyword match is enough per video
status = "COVERED_ON_CHANNEL" if matching else "NOT_COVERE
…
## Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- **Author:** [EricTechPro](https://github.com/EricTechPro)
- **Source:** [EricTechPro/startup-claude-skills](https://github.com/EricTechPro/startup-claude-skills)
- **License:** MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.