AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Docs Crawler

skill-caesiumy-ko-design-md-docs-crawler · by CaesiumY

>-

No reviews yet
0 installs
20 views
0.0% view→install

Install

$ agentstack add skill-caesiumy-ko-design-md-docs-crawler

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-caesiumy-ko-design-md-docs-crawler)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Docs Crawler? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

docs-crawler

Crawl a documentation website into a single Markdown corpus that an LLM can consume as context.

When to use

The user wants to capture a multi-page documentation, design-system, or reference site as Markdown — to feed an LLM, archive it, or research a design system. This skill is also invoked by the design-md skill's research phase.

Inputs

  • Site URL (required) — the documentation site, e.g.

https://socarframe.socar.kr/. Any page on the site works; page discovery starts from the site's sitemap.xml.

  • Output directory (optional) — where to write the corpus. Defaults to

./{host}-docs/ in the current working directory.

If the user has not provided a URL, ask for one — do not guess.

Running the crawl

From the repository root:

pnpm crawl:docs  [--out ] [--external-images]

Equivalent direct form:

pnpm exec tsx .claude/skills/docs-crawler/scripts/crawl.ts  [--out ]

The engine discovers pages via sitemap.xml (falling back to same-origin link-following), fetches each page, extracts the main content, and converts it to Markdown. It prints per-page progress and a final summary line.

By default it also localizes images: every external image and inline base64 data: image is downloaded into crawl/images/ and the Markdown is rewritten to relative paths, so the corpus is self-contained for renderers that can't fetch external URLs (Claude Design, offline previews). Pass --external-images to skip downloading and keep the original external URLs instead.

Output

These artifacts are written under the output directory:

  • crawl-corpus.md — every successful page merged into one document, with

a table of contents and a Source: URL per page. This is the primary deliverable.

  • crawl/pages/{NNN}-{slug}.md — one file per page, each with

source_url / title / method frontmatter.

  • crawl/images/ — downloaded image files (skipped with

--external-images), referenced from the corpus and page files by relative paths.

  • crawl/manifest.json — a per-URL audit log (status, fetch method,

extracted character count, error reason).

After crawling

Report back to the user:

  • How many pages succeeded versus failed (from the summary line or

crawl/manifest.json).

  • The path to crawl-corpus.md.
  • If any pages failed, list them from the manifest with the reason (e.g.

disallowed by robots.txt, no extractable content).

Do not paste the whole corpus into the conversation — point the user at the file. The corpus can be large.

Output location and the design-md skill

The --out argument controls where everything is written. When the design-md skill calls this crawler during its research phase, it passes --out pointing at its cache directory (.claude/cache/design-md/{slug}/), so the resulting crawl-corpus.md (and the localized crawl/images/) land where research-collector can read and cite them. The whole cache directory is gitignored, so downloaded images stay out of version control. Nothing special is needed for that case — just run the crawl with the given --out.

Notes and edge cases

  • No sitemap — discovery automatically falls back to following same-origin

links from the start URL.

  • JavaScript-rendered sites — if a page returns an empty shell, the engine

re-fetches it through a headless browser. Chromium is installed automatically on first need (~150 MB, one-time). Static sites never launch a browser.

  • Images are downloaded into crawl/images/ by default and referenced by

relative paths, so the corpus is self-contained — this covers external ` URLs and inline base64 data: images alike. A non-base64 (percent-encoded) data: URI, or any image whose bytes fail to download or decode, instead becomes an inline-image-omitted placeholder (alt and title text are kept) so raw base64 never bloats the corpus. With --external-images the crawler downloads nothing: external URLs stay as-is and inline data:` images collapse to the placeholder.

  • Few or no images is normal, especially for design-system sites — modern

docs and design-system sites (Docusaurus and similar) render icons and component visuals as inline `, CSS, or live DOM rather than files; any real ` (e.g. a navbar logo) is usually site chrome the content extractor correctly drops. An image-light or image-free corpus is expected in that case — it is not a crawl failure. Capture visuals separately (screenshots) if the research needs them.

  • robots.txt is respected; disallowed paths are skipped.
  • Page cap — at most 200 pages per crawl.
  • Zero pages crawled — the engine exits non-zero and prints the reason

(bad URL, blocked, or an empty site). Surface that to the user rather than reporting success.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.