AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Crawlie

mcp-spronta-crawlie · by spronta

Fast, free, open-source technical SEO + GEO crawler: built for humans and agents.

No reviews yet
0 installs
38 views
0.0% view→install

Install

$ agentstack add mcp-spronta-crawlie

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-spronta-crawlie)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Crawlie? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

crawlie

The fast, free, open-source technical SEO + GEO crawler — built for humans and agents.

Crawl any site for broken links, redirects, missing metadata, and 40+ SEO & Generative-Engine checks — with plain-English guidance on every fix. Runs locally, ships a CLI and an MCP server, and costs nothing.

[](https://www.npmjs.com/package/crawlie) [](https://github.com/spronta/crawlie/actions/workflows/ci.yml) [](LICENSE)

Setup · CLI · MCP & agents · Changelog · Use cases · Why I built this · Desktop app · Checks · Compare · Architecture

Read the docs → crawlie.dev

💡Share feedback & suggestions here

.png)


Setup

The easy way — npm (installs the crawlie CLI and the crawlie-mcp server):

npm i -g crawlie

The macOS app — grab the signed .dmg from Releases.

From source — needs Rust (engine/CLI/MCP) and, for the desktop app, pnpm + Node:

git clone https://github.com/spronta/crawlie
cd crawlie
cargo build --release
# → target/release/crawlie  and  target/release/crawlie-mcp

# or install onto your PATH:
cargo install --path crates/crawlie-cli      # installs `crawlie`
cargo install --path crates/crawlie-mcp      # installs `crawlie-mcp`

> How it ships: the CLI + MCP come only through npm — the right native binary installs automatically as a platform package (nothing to download or unblock). The desktop app is the only direct download: a signed, notarized .dmg on Releases.


How to use (CLI)

# Crawl a whole site (respects robots.txt, seeds from sitemap.xml)
crawlie crawl https://example.com --format pretty

# Audit a single page, or a specific set of pages
crawlie audit https://example.com/pricing
crawlie audit https://example.com/a https://example.com/b

# Save a shareable, self-contained HTML report
crawlie crawl https://example.com --format html -o report.html

# Clean JSON on stdout (perfect for piping / scripting / agents)
crawlie crawl https://example.com --format json -o report.json

# Learn why any finding matters and how to fix it
crawlie explain geo-not-answerable

Output formats: pretty (terminal), json (machine-readable, the default), csv (issues), html (shareable file).

Common flags:

| Flag | What it does | |---|---| | --max-pages | Cap pages fetched (default 500) | | --max-depth | Max click depth from the seed | | --concurrency | Parallel requests (default 16) | | --include / --exclude | Scope the crawl by URL pattern | | --no-robots / --no-sitemap / --no-external | Turn off robots.txt, sitemap seeding, external link checks | | --severity error\|warning\|notice | Only output findings at/above a level | | --save | Save to local report history (crawlie reports, crawlie report ) | | --fail-on error\|warning | Non-zero exit code for CI gating |

Every crawl returns two scores: a Health score (technical SEO) and a GEO score (AI-search readiness).


Use with agents (MCP)

crawlie ships a Model Context Protocol server so an LLM agent can run a full audit and act on it — no human in the loop. This is the part most SEO tools don't have.

Connect it

After npm i -g crawlie, crawlie-mcp is on your PATH. For Claude Desktop, edit claude_desktop_config.json:

{
  "mcpServers": {
    "crawlie": {
      "command": "crawlie-mcp"
    }
  }
}

For Claude Code:

claude mcp add crawlie crawlie-mcp

(If you built from source instead, use the absolute path to target/release/crawlie-mcp.)

(Any MCP-compatible client works — Cursor, Cline, your own agent. It speaks JSON-RPC over stdio.)

One-step install: the Claude Code plugin

The fastest path. The [crawlie plugin](.claude-plugin/plugin.json) bundles the MCP server and a set of skills (audit playbooks) in a single install — the MCP server auto-runs via npx, so you don't even pre-install the binary:

# add this repo as a marketplace, then install the plugin
claude plugin marketplace add spronta/crawlie
claude plugin install crawlie@spronta

Skills (works with any agent, even without the MCP)

The [skills/](skills/) folder holds standalone Agent Skills that teach an agent how to run real audits — full-site SEO + GEO, broken-link fixes, pre-launch gates, and AI-search readiness. Each is self-contained: it needs neither this repo nor a pre-installed crawlie. Missing the binary? The skill runs it on demand via npx -y -p crawlie … (the install is the run), and automatically uses the MCP tools when they're present. See [skills/README.md](skills/README.md).

Tools exposed

| Tool | Purpose | |---|---| | crawl_site | Crawl + audit a whole site (SEO + GEO), returns scores, issues, per-page data | | audit_url | Audit a single page | | audit_urls | Audit an explicit list of pages | | explain_issue | Why a rule matters + how to fix it | | list_rules | The full catalogue of checks | | list_reports / get_report | Read saved crawl history |

Example agent prompts

> "Crawl crawlie.dev, then give me the top 5 fixes that would most improve my GEO score, with the exact change for each."

> "Audit these three landing pages and tell me which is least ready to be cited by AI search, and why."

> "Run a crawl with --fail-on error semantics — are there any broken links or 5xx pages blocking launch?"

The agent calls crawl_site, reads the structured issues, and uses explain_issue to turn findings into a prioritized, actionable plan.


Use cases

  • Pre-launch QA — catch broken links, redirects, 4xx/5xx, and missing metadata before you ship.
  • GEO optimization — make pages citable by AI search: structured data, semantic HTML, answer-ready content, authorship/E-E-A-T.
  • Agent workflows — let a marketing/SEO agent audit a site and propose fixes autonomously via MCP.
  • CI/CD gatingcrawlie crawl … --fail-on error in a pipeline to block regressions.
  • Client reporting — generate a polished, shareable HTML report in one command.
  • Auditing AI-generated sites — verify that the site your agent just built is actually built for search.

Why I built this

I'm Sean Ryan. I've spent 6+ years as a Lead Marketing Engineer, and on the side I build AI tooling for marketers.

With AI, it's faster than ever to ship a marketing site — but most of what gets generated is slop that was never built to be found. And the tools meant to catch that fall short: most SEO auditors cost money, don't play nicely with your agents, or tell you what's wrong without telling you how to actually rank for SEO and GEO (Generative Engine Optimization — being cited by AI search like ChatGPT, Perplexity, and Google AI Overviews).

crawlie fixes that. It's free, it's local-first, it's agent-native, and every issue it finds comes with why it matters and how to fix it.

If this is useful to you, connect with me on LinkedIn → — I share what I'm learning building AI for marketers and SEO/GEO tooling, and I'd love to hear how you're using crawlie.


Desktop app

A beautiful Tauri + React app (Geist design, light/dark, seamless window chrome):

cd apps/desktop
pnpm install
pnpm tauri dev          # live native crawls
pnpm dev                # preview the UI in a browser (demo data, no backend)

Whole-site / single-page / URL-list modes, live progress, Health & GEO score rings, issues with built-in why-it-matters guidance, a sortable pages table, a per-page drawer (GEO signals, headers, schema, hreflang…), auto-saved report history, and one-click shareable HTML export.

> First run, generate the icon set: cd src-tauri/icons && python3 generate.py && cd .. && pnpm tauri icon icons/source.png


What it checks

50 rules and counting.

Technical SEO — broken links · 4xx/5xx · redirects & chains · titles & meta descriptions (missing / duplicate / length) · H1s · canonicals · noindex / nofollow / X-Robots-Tag · robots.txt blocking · images missing alt · thin & duplicate content · orphan & deep pages

Performance & security — slow responses · large pages · missing compression · HTTPS · mixed content · HSTS

Mobile, international & social — viewport · lang · hreflang · Open Graph · Twitter cards · structured data

Structured-data validation — parses JSON-LD and checks each item against Google's rich-result requirements: invalid markup, missing required fields, and missing recommended fields (Product, Article, Recipe, Event, FAQ, Breadcrumb, and more)

JavaScript rendering — crawl with --render to audit each page's post-JavaScript DOM via headless Chrome, so client-rendered content (React/Next/Vue) is seen, and content-requires-js flags pages whose content only exists after JS runs

GEO — Generative Engine Optimization — structured data, semantic HTML, answer-readiness, authorship/E-E-A-T, dated content, question-style headings, and extractable blocks, rolled into a per-page GEO score.

Every finding links to plain-English guidance: why it matters, how to fix it, and what happens if you ignore it.


How it compares

| | crawlie | Screaming Frog | Sitebulb | |---|:---:|:---:|:---:| | Price | Free & open-source | £259/yr to unlock | from £13.50/mo | | Engine | Rust, async, tiny binary | Java (JVM) | .NET | | CLI with JSON output | ✅ | partial | ❌ | | JavaScript rendering | ✅ headless Chrome | ✅ | ✅ | | MCP server (agent-native) | ✅ | ❌ | ❌ | | GEO — AI/answer-engine audit | ✅ | ❌ | ❌ | | "Why it matters" built in | ✅ every issue | ❌ | partial | | Shareable HTML report | ✅ | paid | ✅ | | Source you can read & extend | ✅ | ❌ | ❌ |


Architecture

crates/
  crawlie-core    # the engine — crawl, audit, score, knowledge base, reports
  crawlie-cli     # `crawlie` — JSON / pretty / CSV / HTML output
  crawlie-mcp     # `crawlie-mcp` — Model Context Protocol server (stdio)
apps/
  desktop         # Tauri v2 + React (Geist) desktop app

crawlie-core has zero host dependencies — the same audited engine drops straight into a cloud worker (it already targets wasm32). One engine, every surface, identical results.


Roadmap

  • Cloud workers (shared Rust core) for scheduled/remote crawls
  • JavaScript rendering for SPA-heavy sites
  • Crawl-to-crawl comparison & regression alerts
  • Internal-link graph visualization

License & author

MIT © Spronta Ltd.

Built by Sean Ryan — Lead Marketing Engineer at Pendo.io, building AI for marketers on the side. Connect on LinkedIn →

If crawlie saves you time, a ⭐ on the repo and a hello on LinkedIn mean a lot.

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.