Install
$ agentstack add skill-qinghonglin-data2story-skill-scout ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ● Environment & secrets Used
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Scout
> Premium-profile stage. The orchestrator runs the Scout only in the premium profile; the fast profile skips it. The "always runs / mandatory BGM, no exemption" rules below apply within premium.
Your job is rich media + freshness, with proof. The Detective already gathered background context and basic reference photos; you go further — you find the emotionally strong media (the star player, the packed stadium), the music that sets the mood, and the latest real-world status — and you are the pipeline's media verifier: every asset you pass on has a checked license that permits republication and a checked identity (it really is what the caption says).
You do not generate media — that is the Designer's job. You find real, license-clean media and you prove it.
Setup
DATA_DIR= first argumentPROJECT_DIR= second argumentSKILL_DIR= the directory containing thisSKILL.md(.../skills/data2story-pro/scout)- Read
PROJECT_DIR/detective.json— itsitemsgive you the subjects/topics; itsreference_media+instancestell you what's already covered, so you don't duplicate. - Read any existing manifests in
PROJECT_DIR/assets/(wikimedia_manifest.json,flags_manifest.json,logos_manifest.json) for the same reason. - You may reuse the Detective's fetchers:
python3 SKILL_DIR/../detective/scripts/fetch_images.py(andfetch_flags.py,fetch_logos.py,fetch_openverse.py). - Output:
PROJECT_DIR/scout.json(write incrementally). Assets →PROJECT_DIR/assets/scout_*(prefixscout_to distinguish from the Detective'sref_*).
When to run (always — the cinematic + BGM are mandatory on EVERY blog)
The Cinematographer scroll background and the front BGM are MANDATORY pipeline stages on every blog — there is no "off" / opt-out, and BGM has no exemption (not even privacy) — so the Scout always runs and always sources a real-image set + a fitting real track, on every topic. Key the flavour off the shared [topic_profile](../references/topicprofile.json) (the S3 classifier the Detective resolved into detective.json; if detective.json carries no resolved topic_profile, the Scout MUST write one into scout.json itself — explicit is_visual + is_computational booleans — because an absent profile is now a hard contract error (topic_profile_unresolved), so it cannot be left unresolved): when is_visual is true (any of visualsubject / event / sport / culture / place / emotional) you source the obvious strong subject photos; when the classifier marked the topic non-visual (abstract, text-only, statistical — economics, elections, public-health stats, finance), you still source a relevant real-image set — historical / archival / atmospheric real photos of the era and subject (for an industrial-revolution / economics story: real factory, loom, worker, steam-engine, trading-floor photos from Wikimedia Commons / public domain). Every topic gets a real-image set for the cinematic backing and a fitting real BGM. The only IMAGE exception is privacy_sensitive: there you do not source real-person imagery even if other visual tags are set (lean on non-person archival / atmospheric photos for the backing). A privacy-sensitive topic still gets a BGM — pick a quiet, non-intrusive, mood-appropriate real track (a restrained classical recording fits well); BGM is mandatory on every blog with no audio.used=false escape.
Step 1 — Music (a REAL sourced track + its cover, never AI) — MANDATORY ON EVERY BLOG
BGM is MANDATORY on EVERY blog with NO exemption: every blog opens with a fitting real-sourced track — there is no audio.used=false and no "skip audio for a sober / abstract / privacy topic" branch. Your job is not to decide whether there is a soundtrack but to source the track whose mood fits this story's tone. Match the mood word to the tone: a sober / computational story (economics, elections, public-health stats, finance) wants a pensive / ambient / minimal / orchestral / nocturne track, not a generic upbeat loop; a celebratory / sport / event story wants epic / anthem / fanfare; a somber story wants elegy / adagio / requiem with no especially strong emotion (quiet, non-triumphant). A fitting restrained track on a sober topic is the right BGM — it is NOT "tonally-wrong filler" to set a quiet ambient bed under a numbers story. You ALWAYS return a license-clean BGM track — if no topic-fitting real track exists, you fall to the classical-recording fallback (rung C below), which always yields a license-clean recording. You never leave a blog without a BGM.
The BGM is a real audio track presented in a self-hosted cover-art card at the TOP of the article, directly below the title, that starts on the reader's first gesture — the album/track art is a spinning vinyl disc (a circular cover that rotates only while playing). The track you source here IS the BGM that plays; the Designer never AI-composes a BGM. (text2music is SFX only — atmospheric sound-design beds for an un-findable sound — and is handled by the Designer, never as the front BGM.) FIT FIRST: source the most recognizable best-FIT real track; license-tier is only a tiebreak among comparably-fitting tracks. If the story HAS a signature anthem/track — an official anthem, the artist the post profiles, the song the story is about — source THAT (rung 2) rather than a generic unrelated mood loop, even though the signature track is the demo-gated rung; a recognizable signature track beats a clean-but-unrelated CC0 loop. Only when no signature track fits the story do you reach for a clean-but-generic mood track (rung 1) or, failing that, the classical floor (rung C). Walk this BGM ladder by fit (not blindly top-down), and never AI-compose the BGM:
- License-clean real track (publishable) — find a freely-licensed instrumental that fits the story's mood / place / era (CC0 / CC-BY / public-domain / explicitly royalty-free), self-hosted so it can ship publicly. Use list → pick → download, not blind first-result:
- List candidates:
python3 SKILL_DIR/scripts/fetch_music.py --list --query --limit 8prints a JSON array (each withid,title,license,spdx,duration_s,source_url) to stdout. Commons audio is sparse — search with a SINGLE broad mood word matched to the story's actual tone (epic,anthem,fanfarefor a triumphant story;pensive,ambient,minimal,orchestral,nocturnefor a sober/analytical one;elegy,adagiofor a somber one); multi-word queries usually return nothing (the script auto-falls-back to single words, but a broad word is more reliable). Pick the mood word from the STORY's emotion, not a default-celebratory one — a flat economics/elections/health-stats story wants a restrained pensive/ambient track (still a real BGM, never "no BGM"). - Pick the best fit: prefer one whose
spdxis on [references/license_allowlist.json](references/license_allowlist.json) and whoseduration_ssuits a loopable BGM. Do not just take #1. - Download your choice by its
id:python3 SKILL_DIR/scripts/fetch_music.py --query --outdir PROJECT_DIR/assets --download scout_bgm --id "". The fetcher also downloads the Commons file's cover-art thumbnail alongside the audio (a*_cover.next to the track, recorded ascover_path/cover_source_urlinmusic_manifest.json) so the now-playing card has a license-clean square cover. If a Commons file has no usable thumbnail, fetch a representative license-clean image for the card viafetch_stock.py/ Commons (Step 3), or leave a designed CSS cover to the Designer (NOT an AI image).
Record the full license + attribution (track and cover). This is the track that actually plays, registered as a Scout sct_ audio item (license-clean → it passes the validate.py license-allowlist gate).
- Copyrighted best-fit real track — self-host for the DEMO, publish-gated — when the song the story is about is itself the right BGM (e.g. an official anthem, the artist the post profiles) and no license-clean track fits as well, self-host that real track + its cover for the demo rather than AI-composing one. Fetch the track + a representative cover from its source (a small
--track-url/--cover-urlhelper onfetch_music.py, or grab them by hand — the World Cup anthem + cover were grabbed manually), then record it with an explicit publish-gate so it is never silently treated as clean:
license.spdx = "All Rights Reserved — demo-only",license.permits_republication = false, and a realsource_url(where the track came from).- It is registered as a Designer
des_audio asset withpublish_blocker: true— NOT as a cleansct_item — so it does not pass thevalidate.pylicense-allowlist gate as clean. Apublish_noteis MANDATORY (the swap target — the clean track or embed to switch to before publishing):validate.pySection 8 hard-errors a gated asset with no swap target, so hand the Designer thepublish_notealong with the track + cover + the gate fields. Note in yourscout.json(e.g. alive_status/note item or the relevantsct_notes) that the BGM is the copyrighted demo track to be registered as ades_publish-blocker. The Auditor raises an advisory publish-blocker and the Programmer renders a "demo-only — must license or swap before publishing" credit line; the demo build is flagged, never blocked. - This rung is the right choice for a story with a recognizable signature track (fit beats license-tier). Fall to rung 1 only when no signature track fits the story and a license-clean track does (then publishable beats gated — a tiebreak among comparably-fitting tracks).
- Embed the official player — if you can neither find a license-clean track nor self-host the copyrighted one, surface the real song as an oEmbed-verified
embed(the official Spotify/YouTube player carries its own rights). For anembed: put the /embed/ player URL inembed_urland the watch/track URL you oEmbed-verified insource_url; setidentity.method="oembed",identity.verified=true, andlicense.permits_republication=false(you are not re-hosting — the platform player carries the license;license.spdxmay be"All Rights Reserved") per [../detective/references/instance_verification.json](../detective/references/instance_verification.json). Thevalidate.pylicense gate skips embeds. An embed does NOT replace a self-hosted now-playing card if rung 1 or 2 was available.
Classical-recording fallback ladder — the GUARANTEED license-clean floor (rung C). When no topic-fitting real track (rung 1) and no signature track (rung 2/3) lands, you do not stop with no BGM — you source a license-clean classical RECORDING. The key correctness point: a public-domain composition (Beethoven / Bach / Chopin / Tchaikovsky / Mozart / Haydn / Brahms / Debussy / Satie…) is **NOT automatically a public-domain recording — the score may be PD while a modern performance is fully copyrighted. So you must source a license-clean RECORDING of the piece and verify the recording's own license**, from a PD/CC recording library:
- Sources for clean recordings: Musopen (PD / CC performances), Wikimedia Commons (PD/CC audio), IMSLP (recordings tab — check each recording's license, not just the score's), Free Music Archive (CC tracks). List → pick → download with the same
fetch_music.py --list … --download …flow; record the recording'sspdx(must be on [references/license_allowlist.json](references/license_allowlist.json)),permits_republication, andattribution_text. Verify the recording (not the composition) is what passes the gate. - Pick the piece by era + mood: prefer a period-appropriate piece (match the topic's era if findable — a 1920s story → a 1920s-era composition; a Renaissance topic → early/Baroque), else a famous master. Keep it mood-appropriate: a somber / sober topic gets a quiet, non-triumphant piece with no especially strong emotion (a nocturne, an adagio, the Gymnopédies, a slow movement), never a triumphant fanfare; a celebratory topic may take a brighter classical piece. The classical floor is REAL recordings — it is never AI-composed.
- Register the chosen classical recording as a clean
sct_audio item (license-clean → it passes thevalidate.pylicense-allowlist gate), with its cover (the album/portrait art the library or Commons provides, else a representative license-clean image for the disc, else a designed CSS disc — never AI). This rung always succeeds, so every blog ends with a license-clean BGM.
You may also record the real songs the story references (an anthem, a viral hit) as oEmbed embed instances for a "listen ↗" link in context even when the BGM is a rung-1 / rung-C track — that is separate from the BGM itself.
License gate: never pass a copyrighted commercial track off as a license-clean sct_ BGM. A copyrighted self-hosted BGM is only the rung-2 des_ publish-blocker path above (flagged, demo-only); a clean sct_ BGM is rung 1 or the rung-C classical recording. A PD composition with a copyrighted recording is NOT clean — verify the recording's license, and if the only available recording is copyrighted, treat it like any copyrighted track (rung 2 demo-gate or rung 3 embed), then keep climbing toward a clean classical recording so the blog ends license-clean.
Weight note: Commons audio is often a multi-MB WAV/FLAC. Pass it on as-is (don't degrade the source), but the Designer will transcode it to a web-weight streaming copy (~128 kbps mp3/opus, : N updates, latest = …"), nothing more. (Shared with the Editor/Designer work-streams; topic-agnostic.)
Step 3 — High-value real media (find better than the Detective got)
For the subjects that carry the story emotionally (named people, specific stadiums / places, key objects), fetch a strong, specific real photo / video the Detective missed or got only weakly. You have three complementary image sources — use whichever lands the better, more specific shot, and you may try more than one:
- Wikimedia Commons (by Wikidata QID) — trusted provenance, best for an entity that has a Wikidata page. Fetch with the Detective's helper using a scout prefix:
python3 SKILL_DIR/../detective/scripts/fetch_images.py --qids --props P18 --outdir PROJECT_DIR/assets --prefix scout_ --append(find the subject's Wikidata QID;P18is the entity's photo). Writesassets/scout_*directly. - Openverse (by keyword) — aggregates Flickr-CC, museums (Met, Smithsonian), Wikimedia and more, so it reaches subjects Commons indexes poorly. List then pick then download:
python3 SKILL_DIR/../detective/scripts/fetch_openverse.py --list --q "" --limit 8returns JSON candidates (each withid,spdx,permits_republication,attribution_text,license_url,foreign_landing_url,source_url); pick one whosespdxis on the allowlist (permits_republication: true), then... --download --id --q "" --outdir PROJECT_DIR/assets --prefix scout_. - Stock — Unsplash / Pexels (by keyword) — free-commercial-use, no-attribution stock with
Unsplash-License/Pexels-License(both on the allowlist, genuinely re-hostable); best for atmospheric / generic / cinematic-background shots (a floodlit stadium, a city skyline, an empty arena) where Commons/Openverse are thin — this is the channel the gold blog's cinematic backdrops drew on. Same list → pick → download:python3 SKILL_DIR/scripts/fetch_stock.py --list --q "" --limit 8 --source bothreturns S2-shaped candidates; pick one, then... --download --id --q "" --outdir PROJECT_DIR/assets --prefix scout_. Needs a free key —UNSPLASH_ACCESS_KEYand/orPEXELS_API_KEY(same env
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: QinghongLin
- Source: QinghongLin/data2story-skill
- License: MIT
- Homepage: https://data2story.github.io/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.