Install
$ agentstack add mcp-xberg-io-crawlberg ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Crawlberg
High-performance Rust web crawling engine for structured data extraction. Scrape, crawl, and map websites with native bindings for 14 languages — same engine, identical results across every runtime.
What and Why?
Crawlberg is the crawling substrate: everything you need to scrape and crawl a site end-to-end from a single Rust core — HTML→Markdown, headless-Chrome fallback, robots/sitemap parsing, per-domain throttling, and an SSRF-safe policy — with identical results across 14 language bindings.
Productization concerns (managed proxy pools, tuned WAF fingerprints, authenticated-session injection, scheduling, billing) live in xberg-enterprise, the reference operational implementation. Every extension point (Frontier, RateLimiter, CrawlStore, EventEmitter, ContentFilter, WafClassifier, …) is a trait you inject via CrawlEngineBuilder::with_(...).
Features
| Feature | Description | | ------- | ----------- | | Structured extraction | Text, metadata, links, images, assets, JSON-LD, Open Graph, hreflang, favicons, headings, response headers | | Markdown conversion | Clean Markdown with citations, document structure, and fit-content mode | | Concurrent crawling | Depth-first, breadth-first, or best-first traversal with configurable depth, page limits, and concurrency | | 14 language bindings | Rust, Python, Node.js, Ruby, Go, Java, Kotlin (Android), C#, PHP, Elixir, Dart, Swift, Zig, and WebAssembly | | Smart filtering | BM25 relevance scoring, URL include/exclude patterns, robots.txt compliance, sitemap discovery | | Browser rendering | Optional headless browser for JavaScript-heavy SPAs with WAF detection and bypass | | Batch & streaming | Scrape or crawl hundreds of URLs concurrently; real-time crawl events via async streams | | SSRF-safe by default | Refuses loopback, private, link-local, and cloud-metadata addresses; opt out via env var or CrawlConfig | | Auth & rate limiting | HTTP Basic, Bearer, and custom-header auth with cookie jars; per-domain request throttling | | MCP server & REST API | Model Context Protocol integration for AI agents plus an HTTP server with OpenAPI spec |
Supported Platforms
Precompiled binaries for Linux (x8664/aarch64), macOS (ARM64), and Windows (x64) across every binding. See the platform support reference for the full matrix.
⭐ Star this repo to show your support — it helps others discover Crawlberg.
Quick Start
Language Packages
Python
pip install crawlberg
See Python README for full documentation.
Node.js
npm install @xberg-io/crawlberg
See Node.js README for full documentation.
Rust
cargo add crawlberg
See Rust README for full documentation.
Go
go get github.com/xberg-io/crawlberg/packages/go
See Go README for full documentation.
Java
Available on Maven Central as io.xberg.crawlberg:crawlberg. See Java README for the dependency snippet and current version.
C#
dotnet add package Crawlberg
See C# README for full documentation.
Ruby
gem install crawlberg
See Ruby README for full documentation.
PHP
composer require xberg-io/crawlberg
See PHP README for full documentation.
Elixir
Add {:crawlberg, "~> 0.3"} to your mix.exs dependencies. See Elixir README for full documentation.
Dart / Flutter
dart pub add crawlberg
See Dart README for full documentation.
Kotlin (Android)
Available on Maven Central as io.xberg.crawlberg.android:crawlberg-android. See Kotlin README for the dependency snippet and current version.
Swift
Add via Swift Package Manager. See Swift README for full documentation.
Zig
See Zig README for installation and usage.
WebAssembly
npm install @xberg-io/crawlberg-wasm
See WebAssembly README for full documentation.
C/C++ (FFI)
C header + shared library from GitHub Releases. See FFI crate for full documentation.
CLI
cargo install crawlberg-cli
brew install xberg-io/tap/crawlberg
See CLI README for full documentation.
AI Coding Assistants
Install the Crawlberg plugin from the xberg-io/plugins marketplace. It ships the Crawlberg agent skills (site crawling, HTML→Markdown scraping, headless-Chrome fallback) plus the crawlberg MCP server, and works with every major coding agent — expand your harness below.
Claude Code
/plugin marketplace add xberg-io/plugins
/plugin install crawlberg@xberg
Codex CLI
/plugins add https://github.com/xberg-io/plugins
Then search for crawlberg and select Install Plugin.
Cursor
Settings → Plugins → Add from URL → https://github.com/xberg-io/plugins, then select crawlberg.
Gemini CLI
gemini extensions install https://github.com/xberg-io/plugins
Factory Droid
droid plugin marketplace add https://github.com/xberg-io/plugins
droid plugin install crawlberg@xberg
GitHub Copilot CLI
copilot plugin marketplace add https://github.com/xberg-io/plugins
copilot plugin install crawlberg@xberg
opencode
Add the package to opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["@xberg-io/opencode-crawlberg"]
}
Documentation
Full guides, per-language API references, the substrate/operational model, antibot strategy, and observability live at docs.crawlberg.xberg.io.
Contributing
Contributions are welcome! See our Contributing Guide.
Part of Xberg.dev
- Xberg — document intelligence: text, tables, metadata from 91+ formats with optional OCR.
- Xberg Enterprise — managed extraction API with SDKs, dashboards, and observability.
- crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
- html-to-markdown — fast, lossless HTML→Markdown engine.
- liter-llm — universal LLM API client with native bindings for 14 languages and 143 providers.
- tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
- alef — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.
License
Links
- Documentation
- GitHub Repository
- Issue Tracker
- [Changelog](CHANGELOG.md)
- Discord — community, roadmap, announcements.
Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: xberg-io
- Source: xberg-io/crawlberg
- License: MIT
- Homepage: https://docs.crawlberg.xberg.io
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.