AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Cockroach Crawler

mcp-ajnasnb-cockroach-crawler · by AjnasNB

Give AI agents bounded eyes on the public web: crawl, search, and normalize evidence without an unrestricted browser.

No reviews yet
0 installs
0 views
view→install

Install

$ agentstack add mcp-ajnasnb-cockroach-crawler

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-ajnasnb-cockroach-crawler)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
yesterday

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Cockroach Crawler? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

[](https://www.npmjs.com/package/cockroach-crawler) [](https://github.com/AjnasNB/cockroach-crawler/releases/tag/v0.7.0) [](https://github.com/AjnasNB/cockroach-crawler/actions/workflows/ci.yml) [](https://nodejs.org/) [](./LICENSE)

Documentation · White paper · [Quickstart](./docs/QUICKSTART.md) · Benchmarks · npm


Measured, not asserted

The stable 0.7.0 release defines separate core, quality, and fail-closed extraction profiles. Every value below is development evidence from the same source-pinned WCEB v1.0 scorer, not held-out confirmation or a universal performance claim.

Install the stable package:

npm install cockroach-crawler@0.7.0

| Surface and corpus | Precision | Recall | F1 | Required-snippet | Unwanted | Abstentions | | --- | ---: | ---: | ---: | ---: | ---: | ---: | | Node quality balanced, observed 511 | 0.894101 | 0.926022 | 0.890524 | 0.864090 | 0.111383 | - | | Node quality balanced, WCEB development 1,497 | 0.852784 | 0.896259 | 0.847064 | 0.755867 | 0.096181 | - | | Node quality balanced + fail-closed, observed 511 | 0.847901 | 0.875080 | 0.844935 | 0.812035 | 0.104207 | 43 |

The upstream dataset names the 511-page partition test, but this project previously inspected and iterated against it. We therefore label it observed development evidence, not an untouched held-out result. The 1,497-page row is the upstream WCEB development split. These results do not support a universal 0.90 precision claim.

Published evidence receipt: the immutable benchmark artifact [wceb-quality-observed-0.7.0.json](./bench/results/wceb-quality-observed-0.7.0.json) unchanged from benchmark-generation commit 90825063d447f07345388d040b1428a311109c2b (research-crawler-0.7.0-evidence-source). The JSON records source version 0.7.0; it was first packaged byte-for-byte in cockroach-crawler@0.7.0-rc.1 at commit 62f270636a019c9bcc617a13fe254640bcd06925. That package commit has a valid GitHub signature. v0.7.0-rc.1 is an annotated tag without a cryptographic tag signature. The result artifact has SHA-256 a71c884e9521d1cd1c6326dc07c1d1a5c36344244c45d4900a078ae92a8de535, pins WCEB v1.0 revision 62ff86d12ea72c80c31fb810ff1a724fad687bea, and uses macro averages of page-level Unicode-word precision, recall, and F1. All 511 pages were accepted by this non-fail-closed profile.

This published result remains valid within that observed-development scope. A separate later raw-DOM experiment was rejected after five gates failed; it did not replace, retract, or modify the benchmark shipped in 0.7.0-rc.1 and carried unchanged into 0.7.0. The maintainer selected stable API and package availability independently of that rejected experiment. This is not a benchmark-leadership claim.

The opt-in cockroach-crawler/quality surface uses the exact native trafilatura@0.2.0 dependency; it is not presented as a new extraction algorithm invented here. Fail-closed mode is a separate safety profile: its 43 abstentions return no body when shell, challenge, size, quality, or output limits make admission unsafe. Method, raw rows, profile definitions, historical Python baseline comparison, and claim boundaries are in [docs/BENCHMARK.md](./docs/BENCHMARK.md) and [docs/EXTRACTION-COMPARISON.md](./docs/EXTRACTION-COMPARISON.md).


Give your AI agents the web. Keep the keys. Crawl complete sites, render JavaScript, rank fetch-validated maps, extract structured data with CSS, XPath, restricted regex, or a schema-validated host model, parse PDFs, and return evidence without handing a model an unrestricted browser or network client.

npm install cockroach-crawler

Cockroach Crawler is an open-source AI web crawler and TypeScript toolkit for agents, RAG pipelines, documentation indexing, research, content inventory, and QA. It turns explicit public URLs and supported read-only sources into clean Markdown, JSON, and JSONL evidence records. Use one package to run BFS, DFS, best-first, or adaptive traversal, build searchable fetch-validated site maps, extract bounded CSS, XPath, or restricted-regex fields, search YouTube without a developer API key through the optional reviewed route, render JavaScript pages, run bounded self-hosted jobs, and preserve source identity, redirects, hashes, warnings, artifacts, and provenance.

It is designed for governed agents that need browser rendering, structured extraction, source evidence, and explicit network authority in one Node.js package. That is the product focus, not a claim of universal superiority or access-control bypass.

Every capability stays behind creator-owned origin, request, byte, redirect, concurrency, and time limits. Use the hardened local crawler for bounded public-web collection, the source router for explicit provider capabilities, optional reach providers for reviewed no-developer-key or session-backed reads, and the restricted self-hosted Worker only for allowlisted sites you operate or trust.

Why developers choose Cockroach Crawler

  • One install, several evidence paths: crawl, map, render, extract, inspect provider availability, and export Markdown, JSON, or JSONL.
  • Deep-crawl strategies that fit the task: BFS, DFS, best-first relevance, or adaptive priority under the same hard crawl budgets.
  • Browser and document evidence: bounded screenshots, PDF generation and parsing, open Shadow DOM and readable iframe flattening, virtual-scroll helpers, and explicit persistent profiles.
  • Deploy it where agents already run: local API, CLI, native MCP server, authenticated Docker API, and browser playground.
  • Agent-ready without ambient authority: model input can narrow a crawl but cannot expand creator-owned origins, private-network access, browser authority, or resource ceilings.
  • Local-first and open source: normal public crawling needs no hosted account or crawler API key.
  • Proof travels with the content: canonical URLs, redirect history, content hashes, retrieval metadata, failures, warnings, and provenance stay attached to results.
  • No-key routes are explicit: public GitHub reads and an optional reviewed YouTube route work without developer API credentials; session-backed providers remain separately installed and operator controlled.

Cockroach Crawler is not a hosted proxy fleet or an access-control bypass. Compared with broad crawling platforms, its differentiator is the inspectable boundary around every returned record. Read the category-based alternatives guide covering Firecrawl, Crawl4AI, Crawlee, Scrapy, Trafilatura, Playwright, Puppeteer, Apify, and ScrapingBee before choosing a crawler.

50 top-level shipped capabilities

This curated top-level catalog makes the package easier to scan; it is not an exhaustive count of every exported function, option, or output field. Every listed item has a public API, command, output contract, test, or dedicated documentation page in stable 0.7.0.

Crawl and discover - 15

  1. Static HTTP crawling
  2. Multiple seeds
  3. Breadth-first traversal
  4. Depth-first traversal
  5. Best-first traversal
  6. Adaptive relevance traversal
  7. Sitemap discovery
  8. Robots enforcement
  9. Include and exclude filters
  10. Validated redirects
  11. Concurrency and politeness
  12. Deadlines and cancellation
  13. Persistent cache
  14. Compact fetch-validated site maps
  15. Searchable fetch-validated site maps

Render and capture - 9

  1. JavaScript rendering
  2. Selector waits and bounded clicks
  3. Infinite and virtual scroll
  4. Open Shadow DOM flattening
  5. Readable same-origin iframe flattening
  6. Full-page screenshots
  7. PDF generation
  8. Trusted operator page hooks
  9. Explicit persistent browser profiles

Extract agent-ready data - 8

  1. Readable Markdown through the dependency-light core or opt-in Node quality backend
  2. CSS schema extraction
  3. XPath extraction
  4. Restricted regex extraction
  5. Optional host-model JSON Schema extraction
  6. Local PDF parsing
  7. Links and page metadata
  8. Evidence hashes and retrieval provenance

Reach public sources - 6

  1. Public GitHub repository and issue reads
  2. YouTube search and metadata without a developer API key through the optional reviewed route
  3. Official YouTube, X, and Reddit provider adapters
  4. Optional read-only session providers for X, Reddit, Facebook, Instagram, LinkedIn, and Xiaohongshu
  5. Offline RSS and Atom parsing
  6. Provider doctor, capability reporting, and deterministic routing

Connect agents - 3

  1. Strict creator-bounded agent tool
  2. Native MCP stdio server
  3. Optional Maqam policy, approval, trace, and evidence integration

Deploy and operate - 4

  1. Authenticated Node.js and Docker API
  2. Responsive dashboard and browser playground
  3. Bounded process-local asynchronous jobs
  4. Fixed-origin Cloudflare Worker profile

Keep authority bounded - 5

  1. Public-network admission and SSRF defenses
  2. DNS pinning and explicit origin policy
  3. Exact resource ceilings
  4. Fixed self-hosted proxy-gateway adapter
  5. Challenge-aware provider escalation that stops without access-control bypass

Open the searchable catalog of 50 top-level capabilities for copyable APIs, expected outputs, prerequisites, failure modes, and the authority boundary for every item.

Web superpowers, bounded by you

  • Turn an explicitly allowed public site or documentation section into normalized text, Markdown, links, metadata, and content hashes under hard page, byte, depth, redirect, concurrency, and time budgets.
  • Read public GitHub repository and issue data through a normalized, read-only adapter that produces the same evidence shape as a crawl.
  • Search YouTube without a developer API key through the optional reviewed yt-dlp route, or use the official API when configured.
  • Read X, Reddit, Facebook, Instagram, LinkedIn, or Xiaohongshu only through an explicitly selected official or operator-controlled read-only route.
  • See JavaScript-rendered pages through an opt-in bounded Chromium mode and expose structural browser actions through a host contract that Maqam can govern.
  • Feed research agents, RAG pipelines, content inventories, QA systems, and audit trails with Markdown, JSON, or JSONL records carrying source identity, retrieval metadata, warnings, hashes, redirects, and provenance.

It does not extract cookies, reuse hidden credentials, bypass logins, CAPTCHA, paywalls, robots policy, or access controls, and it exposes no social write actions.

Public and authenticated GitHub reads

Public repository search, repository reads, and issue reads work without a GitHub token at the public REST limit. Supplying GITHUB_TOKEN or GH_TOKEN raises the documented GitHub API limit while keeping the adapter read-only:

GITHUB_TOKEN=your_token cockroach-sources search github \
  "topic:web-crawler language:javascript" \
  --max-results 10 \
  --json

Tokens are accepted only through trusted configuration or environment state. They are never accepted as a CLI flag and never appear in normalized records, errors, URLs, doctor output, or provenance.

This is separate from MCP Registry ownership. The package carries the case-sensitive registry identity io.github.AjnasNB/cockroach-crawler; GitHub-authenticated registry publication proves control of that namespace, not additional crawler runtime authority.

Version 0.7.0 promotes the reviewed RC runtime represented by commit 62f270636a019c9bcc617a13fe254640bcd06925 without runtime, dependency, or benchmark drift. Stable publication is bound to the exact reviewed 0.7.0 commit and tarball through npm trusted publishing. Verify the live registry state with npm view cockroach-crawler version dist-tags.

Public benchmark evidence

The source-pinned quality balanced profile produced 0.894101 precision, 0.926022 recall, and 0.890524 macro word F1 on the observed 511-page partition. On the 1,497-page WCEB development split it produced 0.852784 precision, 0.896259 recall, and 0.847064 F1.

Separate public-source conformance probes passed 25/25 adapted Google robots.txt dispatch vectors and 101/101 applicable credential-free HTTP(S) canonicalization cases from the pinned Web Platform Tests URL corpus.

The 511-page result is observed development evidence because it influenced earlier iterations; it is not an untouched final test. These are reproducible workload results, not a universal quality, speed, RFC certification, or competitor-ranking claim. Read the [method and complete profile table](./docs/BENCHMARK.md), inspect the [machine-readable results](./bench/results/), and run the [source-pinned evaluators](./bench/public/). The local 120-page throughput fixture remains separate because extraction quality and loopback speed measure different things.

The local crawler produces structured JSON/JSONL with readable text, Markdown, links, response metadata, redirect provenance, and content hashes for documentation indexing, RAG ingestion, content inventory, QA, research, and agent tools. The source adapters normalize GitHub, YouTube, X, and Reddit records when each provider's documented access requirements are met.

It does not bundle a model, model key, hosted account, stealth layer, CAPTCHA bypass, paywall bypass, or authorization bypass. Optional LLM extraction runs only through a host-supplied adapter and validates its output against the caller's JSON Schema.

Documentation: quickstart · 50-capability top-level catalog · [advanced capabilities](./docs/ADVANCED.md) · [detailed feature inventory](./docs/FEATURES.md) · comparison · benchmark · [architecture](./docs/ARCHITECTURE.md) · [source adapters](./docs/SOURCES.md) · GitHub Discussions · issues · [security](./SECURITY.md) · [contributing](./CONTRIBUTING.md)

Complete documentation

The website is organized as a task manual plus a typed reference. Every guide uses stable public exports and copyable examples from this package.

| Need | Guide | | --- | --- | | Install and run one bounded crawl | Documentation overview | | Browse the 50 curated top-level capability pages with API, output, failures, and boundaries | Top-level capability catalog | | Follow the CLI quickstart and bounded workflow | CLI guide | | Embed the typed Node.js API | JavaScript API | | Configure BFS, DFS, best-first, adaptive traversal, sitemaps, callbacks, and cache | Crawling and cache | | Render JavaScript, click, scroll, flatten DOM, capture screenshots and PDFs, and use explicit profiles | Browser rendering and evidence | | Generate core or Node-quality Markdown, fail closed on low-confidence pages, or extract with CSS, XPath, restricted regex, local PDF parsing, or a host-supplied model adapter | Extraction manual | | Create compact fetch-validated maps and optionally rank them by search terms | Map and extract | | Give a model

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.