AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Web Extract

skill-serendipityoneinc-zoodata-skills-web-extract · by SerendipityOneInc

>

No reviews yet
0 installs
0 views
view→install

Install

$ agentstack add skill-serendipityoneinc-zoodata-skills-web-extract

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-serendipityoneinc-zoodata-skills-web-extract)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Web Extract? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Web Extract — Structured Data from the Open Web

Backed by the ZooData WebTools API. Six HTTP endpoints. One API key. Structured JSON by default — no second LLM pass to parse fields.

Files

| File | Purpose | |------|---------| | {skill_base_dir}/scripts/webtools.py | Thin CLI wrapper — one subcommand per endpoint. Has the Cloudflare-UA, crawl-warmup, and nested-data-shape quirks baked in. Run --help for params. | | {skill_base_dir}/references/reference.md | Full request/response schemas, error codes, billing, edge cases. Load when you need exact field names. |

Why pick this skill (vs the alternatives in this environment)

| Tool | What it gives you | When to pick it | |---|---|---| | web-extract (this skill) | Page → {title, summary, sections, key_metrics, outgoing_links, ...} JSON in one call | You need page DATA (price, specs, fields, link graphs) — downstream code or LLM can use the JSON directly without re-parsing | | browser-act | Browser session: click, scroll, type, screenshot | You need to interact with a page (login flow, take a screenshot, fill a form) or visually verify rendering | | WebFetch (built-in) | Static URL → markdown | You need a single static page as prose, no JS rendering, no structured fields | | deep-research | Multi-source research with citations | You need a synthesized report drawing from many web sources, not raw data | | monid | Generic tool-discovery layer | You're not sure which tool to use yet and want to browse options |

Key advantage: every other tool above forces a second pass (re-LLM the markdown / re-parse the HTML / extract structure manually). The underlying API returns the structured fields directly — saves a round-trip and tokens.

Credential

Required: ZOODATA_API_KEY (legacy APICLAW_API_KEY still works as fallback). Get a free key (1,000 credits) at zoodata.ai/en/api-keys.

export ZOODATA_API_KEY='hms_live_xxx'
# OR persist to disk:
mkdir -p ~/.zoodata && echo '{"api_key":"hms_live_xxx"}' > ~/.zoodata/config.json

Two ways to drive it

  1. webtools.py CLI (preferred — quirks baked in): UA header, warmup tolerance, retry/backoff, JSON-first defaults all handled. Just run python {skill_base_dir}/scripts/webtools.py .
  2. Raw curl / HTTP (when CLI isn't installed or for ad-hoc calls): every example below also shows the raw POST. Always set User-Agent: web-extract-skill/1.0 — the Cloudflare edge rejects the default Python-urllib UA with HTTP 403 (see Tips).

Endpoint base URL: https://api.zoodata.ai/openapi/v2/webtools/*. All POST with JSON body, except crawl/{id} (GET).

Default format: JSON

This skill defaults to formats: ["json"] for every scrape / scrape-interactive / search-with-scrapeOptions call. The API itself defaults to ["markdown"], but JSON gives the agent structured fields (title, summary, sections, keymetrics, outgoinglinks, page_type, …) that compose better with downstream tools and are usually smaller to put back into context.

Switch to markdown only when the user explicitly asks for "the article text" / "clean prose" / "as markdown", or when the page is genuinely article-shaped (long-form blog post / news / docs) and structured fields wouldn't help. Switch to rawHtml only when the user needs DOM access (custom CSS selector extraction, table parsing the API didn't structure, embedded ` data, etc.). Request multiple formats in one call (["json","markdown"]`) when you genuinely need both — but never silently fall back to markdown when the user didn't ask for it.

When to use

| User intent | Endpoint | Cost | |---|---|---| | Fetch one URL as structured JSON (default) | POST /scrape formats:["json"] | 1 credit | | Fetch one URL as markdown (user explicitly asked) | POST /scrape formats:["markdown"] | 1 credit | | Fetch one URL as raw HTML (user needs DOM) | POST /scrape formats:["rawHtml"] | 1 credit | | Same as above, but page needs JS / login / click-through / scroll | POST /scrape-interactive | 1 credit | | Google search → list of result URLs | POST /search (SERP-only) | 1 credit | | Google search → URLs + structured JSON for each | POST /search scrapeOptions:{format:"json"} | 1 credit per returned result | | List all URLs reachable from a website | POST /map | 1 credit | | Recursively pull every page of a site (async) | POST /crawl + GET /crawl/{id} poll loop | 1/submit + 1/poll |

Quick start (CLI — recommended)

Shorthand: WT="python {skill_base_dir}/scripts/webtools.py"

# 1. Scrape one URL as structured JSON (default)
$WT scrape https://example.com
# → data.json = {title, summary, sections, key_metrics, outgoing_links, author, date, page_type, ...}

# 2. Scrape as markdown (user explicitly asked for prose)
$WT scrape https://example.com --format markdown

# 3. Need BOTH JSON and markdown in one call
$WT scrape https://example.com --format json --format markdown

# 4. JS-heavy page — wait, click "Load more", scrape
$WT interactive https://example.com --actions '[
  {"type":"wait","milliseconds":1500},
  {"type":"click","selector":"button.load-more"},
  {"type":"wait","milliseconds":1000}
]'

# 5. SERP-only search (cheap — 1 credit total, regardless of result count)
$WT search "best wireless earbuds 2026" --limit 10

# 6. Deep-scrape every search result as JSON (1 credit per result)
$WT search "best wireless earbuds 2026" --limit 5 --deep-scrape

# 7. Search restricted to specific domains, last 24h
$WT search "python async" --limit 20 --tbs qdr:d \
  --include-domains "python.org,realpython.com"

# 8. Discover every URL on a docs site under /api/
$WT map https://docs.example.com --include-paths "/api/.*" --limit 1000

# 9. Crawl + poll in one shot (handles warmup, paginates pages internally)
$WT crawl-wait https://blog.example.com --limit 200 --max-depth 3 \
  --poll-interval 10 --max-wait 1800

# 9b. Or submit + manual poll
$WT crawl https://blog.example.com --limit 200 --max-depth 3
$WT crawl-status job_xxx --skip 0 --limit 100

# 10. Verify auth + credits (1 credit — scrapes example.com)
$WT check

Run $WT --help for every flag on any subcommand.

Same calls via raw curl

When the CLI isn't available (no python3 / offline / restricted env), every endpoint also takes a plain JSON POST:

curl -sS -X POST https://api.zoodata.ai/openapi/v2/webtools/scrape \
  -H "Authorization: Bearer $ZOODATA_API_KEY" \
  -H "Content-Type: application/json" \
  -H "User-Agent: web-extract-skill/1.0"   `# REQUIRED — see Tips` \
  -d '{"url":"https://example.com","formats":["json"]}'

See references/reference.md for the full JSON body shape of every endpoint. Any non-curl HTTP client must set a custom User-Agent header — Cloudflare's edge blocks the default Python-urllib UA with HTTP 403 (error code 1010).

Key parameters at a glance

Full schemas are in references/reference.md. Load that when you need exact field names or are debugging.

/scrape: url (required), formats: ["json"|"markdown"|"rawHtml"]skill default ["json"] (API itself defaults to ["markdown"] if formats omitted; always pass it explicitly). Max 3 formats per call.

/scrape-interactive: url, formats, actions: [...] — 7 action types (wait/click/write/press/scroll/scrape/executeJavascript). Max 50 actions, cumulative wait ≤ 60s.

/search: query (supports inline Google operators), limit (1–20, default 10), sources (["web","news","images"]), tbs (qdr:d|w|m|y), includeDomains/excludeDomains (bare hostnames, max 20 each), scrapeOptions: {format:"markdown"|"json"} (presence triggers deep-scrape).

/map: url, limit (1–100000, default 5000), search (keyword rank filter), sitemapMode (include/only/skip), includeSubdomains (default true), includePaths/excludePaths (regex), ignoreQueryParameters (default true).

/crawl: url, limit (1–10000, default 100), maxDepth (1–10), includePaths/excludePaths, allowSubdomains, sitemapMode, ignoreQueryParameters (default false — opposite of /map), crawlEntireDomain.

GET /crawl/{job_id}: query params skip, limit (1–1000). data is an array of page-scrape blobs. meta.total is the running total; compare against data.length to detect completion.

Tips

  • 🚨 Cloudflare 1010 trap (Python / non-curl HTTP clients). The Cloudflare edge in front of api.zoodata.ai blocks default Python Python-urllib/X.Y User-Agent → HTTP 403. ALWAYS set an explicit User-Agent header (any non-default value works, e.g. User-Agent: web-extract-skill/1.0). curl is fine out of the box; Python / Go / Node fetch / requests all need it.
  • Crawl GET /crawl/{job_id} warmup gotcha: for ~5–10 seconds after POST /crawl returns the job_id, polling may return NOT_FOUND. This is the job still registering, not a real "not found." Tolerate it: retry with 5s sleep up to ~3 times before giving up.
  • crawl-status response shape: data is an object {id, status, completed, total, data: [pages...]} — the page array is nested at data.data. Status values: "queued", "running", "completed", "failed". (Failed success:false responses cost 0 credits.)
  • A success:true response can still contain an error page. Always check data.meta.statusCode (or per-result meta.statusCode for search) before trusting content.
  • Deep-scrape /search may return fewer than limit results — pages with no extractable content in the chosen format are dropped.
  • Domain filters are bare hostnames only. https://github.com or github.com/path → 422. Use github.com.
  • includePaths / excludePaths are regex, not glob. Use /blog/.* not /blog/**.
  • Crawl job IDs are tenant-scoped. Polling someone else's job returns 404 (indistinguishable from "job not found" — by design).
  • Refused requests (ACCESS_DENIED, RATE_LIMITED) are NOT billed. TIMEOUT and CONTENT_UNAVAILABLE are.
  • Polling costs 1 credit per call. For large crawls, batch your polls — fetch with limit=1000 and space polls ≥10s apart.
  • Format choice (skill default = JSON): json is the default — gives structured fields (title/summary/sections/key_metrics/outgoing_links/page_type/...) that downstream tools can consume directly. Use markdown only when the user asks for clean prose or the page is article-shaped (blog post / docs); use rawHtml only when DOM access is needed (custom selectors, table parsing the JSON layer doesn't expose).

On Missing Key (no credentials configured)

BEFORE calling any endpoint, verify a credential is configured. Reliable check: python {skill_base_dir}/scripts/webtools.py check — exits 2 if no key is found in env (ZOODATA_API_KEY / legacy APICLAW_API_KEY) or config files (~/.zoodata/config.json / ~/.apiclaw/config.json). A [ -z "$ZOODATA_API_KEY" ] test alone is NOT sufficient.

When no key is found through any mechanism:

  1. STOP. Do not call any endpoint.
  2. Do NOT fall back to "partial output from training data" / "describe the page from common knowledge" / "for reference only" preview. This skill's deliverable is structured extraction of THIS URL's actual content — without the data, there is no deliverable.
  3. Tell the user, in their language, all three of:
  • "ZOODATA_API_KEY is not set — I need this to fetch the page."
  • Get a free key (1,000 credits, no credit card): https://zoodata.ai/en/api-keys
  • Configure via one of:
  • export ZOODATA_API_KEY='hms_live_xxx' (session only)
  • mkdir -p ~/.zoodata && echo '{"api_key":"hms_live_xxx"}' > ~/.zoodata/config.json (persistent)

On 401 Invalid Key

When any endpoint returns HTTP 401:

  1. STOP further calls immediately. A rejected key won't be accepted on retry — every subsequent call will return 401 too.
  2. Tell the user: their ZOODATA_API_KEY was rejected (invalid, revoked, or expired). Direct them to zoodata.ai/en/api-keys.
  3. If you collected partial output before failure, show it and mark partial. Do not fabricate the rest — no "training-data fallback" / "page description from common knowledge" substitution.

On 402 Credit Exhausted

When any endpoint returns HTTP 402:

  1. STOP further calls.
  2. Report partial findings already gathered, plus the creditsRemaining number from the last successful call.
  3. Point the user at zoodata.ai/en/pricing to top up. Do not fabricate the rest — no "common-sense page summary" filler in place of a real fetch.

See also

  • [zoodata](../zoodata/SKILL.md) — commerce endpoints (Amazon products / markets / reviews / brands). Pair with webtools when you need to validate Amazon findings against the open web (competitor sites, news, off-Amazon reviews).
  • [amazon-market-entry-analyzer](../amazon-market-entry-analyzer/SKILL.md) — uses the commerce side; webtools complements it for non-Amazon channel research.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.