AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL unreviewed MIT Self-run

Firecrawl Research Patterns

skill-terrylica-cc-skills-firecrawl-research-patterns · by terrylica

Programmatic Firecrawl usage, self-hosted operations, academic paper routing, recursive deep research, and raw corpus persistence.

— No reviews yet
0 installs
20 views
0.0% view→install

Install

$ agentstack add skill-terrylica-cc-skills-firecrawl-research-patterns

Open-source listing, not yet scanned by AgentStack. Follow the source repository for install instructions.

Security review

⚠ Flagged

1 finding(s); flagged for manual review. · v0.1.0 How review works →

  • • Prompt-injection patterns
  • • Secret / credential exfiltration
  • • Dangerous shell & filesystem operations
  • • Untrusted network calls
  • • Known-malicious package signatures
  • high Reads credentials/environment and may exfiltrate them.

What it can access

  • ● Network access Used
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ● Environment & secrets Used
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Reliability & compatibility

— Not yet reviewed
0 installs to date
— no reviews yet
● 2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Firecrawl Research Patterns? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Firecrawl Research Patterns

Programmatic patterns for using self-hosted Firecrawl in research workflows — search, scrape, route academic papers, run recursive deep research, and persist raw results for future re-analysis. Also covers self-hosted deployment, health checks, and recovery.

For archiving AI chat conversations (ChatGPT/Gemini shares), see Skill(gh-tools:research-archival).


> Self-Evolving Skill: This skill improves through use. If instructions are wrong, parameters drifted, or a workaround was needed — fix this file immediately, don't defer. Only update for real, reproducible issues.

FIRST — TodoWrite Task Templates

MANDATORY: Select and load the appropriate template before any research work.

Intent routing — AI chat share URLs (chatgpt / gemini / claude)

AI chat share URLs (chatgpt.com/share/*, chat.openai.com/share/*, gemini.google.com/share/*, g.co/gemini/share/*, claude.ai/share/*, claude.ai/chat/*) can be processed by either this skill or Skill(gh-tools:research-archival). Pick by intent, not URL pattern:

| Your intent | Skill | Output | | ---------------------------------------------------------------------------------- | ----------------------------------- | ------------------------------------------------------------------------------ | | One-off read / extract conversation text for analysis | This skill — port 3003 (Sec. 5) | Markdown file on Caddy; no frontmatter, no Issue, no provenance. | | Long-term archive with identity verification, frontmatter, GitHub Issue cross-link | Skill(gh-tools:research-archival) | docs/research/YYYY-MM-DD-{slug}-{type}.md + issue with Discovery Provenance. | | Already have the file, just need to scrape extra content into the same corpus file | This skill | Append-mode workflow under your control. |

> Both paths share the same Firecrawl backend. research-archival calls Firecrawl too — it adds an archival layer on top. There is no scraping capability gap between the two; the difference is what happens to the bytes after they come back.

WebFetch limitation, regardless of intent: Claude Code hard-blocks WebFetch against chatgpt.com. Use Firecrawl (this skill, any port) or Jina Reader instead. Verified 2026-05-27.

Empirical note (2026-05-27): port 3003 successfully scrapes ChatGPT shares — curl :3003/scrape?url=...&name=... returned a 75 KB / 1,734-line markdown for a real ChatGPT share via the Caddy two-step pattern (see Section 5). Earlier guidance that said "route AI chat shares out" was overcautious and contradicted Section 5's port table.

Template A — Single Firecrawl Search + Persist

1. Health check — GET http://littleblack.tail0f299b.ts.net:3002/ (expect 200 + {"message":"Firecrawl API",...}; NEVER use /v1/health — it 404s)
2. Execute search — POST /v1/search with query, limit, scrapeOptions
3. Persist raw results — save each result page to docs/research/corpus/ with frontmatter
4. Update corpus index — append entries to docs/research/corpus-index.jsonl
5. Extract findings — summarize key learnings from raw corpus files

Template B — Academic Paper Retrieval + Persist

1. Identify source — classify URL/DOI per academic-paper-routing.md decision tree
2. Route to scraper — arxiv direct HTML, Semantic Scholar API, Firecrawl, or Jina Reader
3. Scrape content — execute fetch with appropriate method and timeout
4. Persist raw result — save to docs/research/corpus/ with academic-specific frontmatter
5. Update corpus index — append entry to corpus-index.jsonl
6. Summarize paper — extract key claims, methods, results from raw corpus file

Template C — Full Recursive Deep Research with Corpus

1. Health check — GET http://littleblack.tail0f299b.ts.net:3002/ (expect 200 + Firecrawl banner; NEVER /v1/health — it 404s)
2. Initialize parameters — set breadth (default 4), depth (default 2), concurrency (default 2)
3. Generate search queries — LLM generates N queries from topic + prior learnings
4. Execute searches — Firecrawl /v1/search for each query via p-limit(concurrency)
5. Persist raw results — save ALL scraped pages to docs/research/corpus/ with provenance
6. Extract learnings — LLM extracts key findings + follow-up questions per result set
7. Recurse — for each follow-up, recurse with breadth=ceil(breadth/2), depth=depth-1
8. Base case — depth=0, return accumulated learnings
9. Synthesize report — LLM generates final markdown from all learnings
10. Write session report — save to docs/research/sessions/ with corpus file references
11. Update corpus index — append all new entries to corpus-index.jsonl

Template D — Corpus Review / Re-Analysis

1. Inventory corpus — read docs/research/corpus-index.jsonl, filter by session/topic/date
2. Read raw files — load matching corpus files from docs/research/corpus/
3. Re-analyze — extract new insights with current context/questions
4. Update session report — amend or create new session report in docs/research/sessions/

Template E — Image-Rich Paper with Inline Figures

Use when paper contains architecture diagrams, result plots, attention maps, or any critical visual content.

1. Scrape text — use port 3003 (preferred, preserves absolute image URLs) or Jina fallback
2. Detect figures — scan scraped markdown for  patterns with .png/.jpg/.svg
3. Extract figure URLs — for arXiv: probe https://arxiv.org/html/{id}v{n}/x{N}.png until 404
4. Keep URLs inline — DO NOT rewrite to local relative paths (breaks GitHub rendering)
5. Ensure inline embedding — markdown body must have  for each figure
6. Catalog in frontmatter — add figure_count and figure_urls list (all absolute URLs)
7. Save corpus file — GFM markdown with inline absolute URLs renders on GitHub without hosting
8. Update corpus-index.jsonl — include has_figures: true, figure_count, figure_urls

Section 1 — Programmatic Firecrawl Usage

Instance: Self-hosted on littleblack — Debian 12 (bookworm), kernel 6.1.0-31, hostname kab, login user yca, RTX 2080 Ti, 62 GiB RAM. No API key required for any Firecrawl endpoint.

| Access path | URL base | When to use | | ------------------ | ------------------------------------------- | -------------------------------------------------------------------------------------------- | | Tailscale FQDN | http://littleblack.tail0f299b.ts.net:3002 | Preferred. Works on every tailnet-attached client regardless of MagicDNS resolver state. | | Tailscale IP | http://100.78.106.112:3002 | Bypasses DNS entirely; stable while the tailnet device exists. | | Tailscale MagicDNS | http://littleblack:3002 | Conditional — only when bare-name resolution works (see preflight below). | | Same-LAN direct | http://192.168.1.67:3002 | Only when the client is on the Telus PureFibre LAN (eno1 interface). | | Legacy ZeroTier | http://172.25.236.1:3002 | Fragile fallback (ztksetviym interface). Prefer Tailscale. |

MagicDNS preflight (run before relying on bare littleblack):

# macOS — does the OS resolver know about the bare name?
dscacheutil -q host -a name littleblack | grep -q '^ip_address'  && echo OK || echo MISSING

# Cross-platform — does any path resolve?
getent hosts littleblack 2>/dev/null || ping -c1 -W1 littleblack 2>&1 | head -1

If preflight returns MISSING / "cannot resolve", use the FQDN row. SSH happens to work because ~/.ssh/config hard-codes the FQDN under the Host littleblack alias — that's an SSH-only shortcut, not a system-wide DNS facility. Bare littleblack over HTTP fails silently as HTTP 000 when the resolver doesn't have it; the failure mode is invisible without ping/dscacheutil. Confirmed broken on m3max (this Mac) as of 2026-05-27.

SSH (for ops, not API calls): ssh littleblack — defined in ~/.ssh/config as HostName littleblack.tail0f299b.ts.net, User yca, IdentityFile ~/.ssh/id_ed25519_zerotier_np.

Why fetch() Instead of @mendable/firecrawl-js SDK

The official SDK uses jiti for dynamic imports, which is incompatible with Bun's module resolution. Direct fetch() calls are simpler, more reliable, and have zero dependencies.

Two Endpoints

| Endpoint | Purpose | When to Use | | ----------------- | --------------------- | ------------------------------------------------- | | POST /v1/search | Search + scrape combo | Research queries — returns multiple scraped pages | | POST /v1/scrape | Single URL scrape | Known URL — extract markdown from one page |

See [api-endpoint-reference.md](./references/api-endpoint-reference.md) for full request/response contracts.

Quick Examples

Use the FQDN base URL — works on every tailnet-attached client regardless of MagicDNS resolver state. Pull from $FIRECRAWL_BASE env var if your project sets one, otherwise hard-code the FQDN:

const FIRECRAWL_BASE =
  process.env.FIRECRAWL_BASE ?? "http://littleblack.tail0f299b.ts.net:3002";

Search (returns multiple results with markdown):

const res = await fetch(`${FIRECRAWL_BASE}/v1/search`, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    query: "mixture of experts scaling laws",
    limit: 5,
    scrapeOptions: { formats: ["markdown"] },
  }),
});
const { data } = await res.json(); // data: [{ url, markdown, metadata }]

Scrape (single URL):

const res = await fetch(`${FIRECRAWL_BASE}/v1/scrape`, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    url: "https://arxiv.org/abs/2401.12345",
    formats: ["markdown"],
    waitFor: 3000, // ms — for JS-heavy pages
  }),
});
const { data } = await res.json(); // data: { markdown, metadata }

Error Handling

// Always set a timeout
const controller = new AbortController();
const timeoutId = setTimeout(() => controller.abort(), 15_000);

try {
  const res = await fetch(url, { ...opts, signal: controller.signal });
  if (!res.ok) throw new Error(`Firecrawl: ${res.status} ${res.statusText}`);
  const json = await res.json();
  if (!json.data || (Array.isArray(json.data) && json.data.length === 0)) {
    // Empty results — not an error, but no content to process
  }
} finally {
  clearTimeout(timeoutId);
}

Health Check

> There is no /v1/health endpoint on this Firecrawl build. Probing it returns HTTP 404 (Express's HTML error page), which looks like a service-down signal but isn't. Use the root / endpoint, which returns HTTP 200 with {"message":"Firecrawl API","documentation_url":"https://docs.firecrawl.dev"}. Confirmed 2026-05-27 against ports 3002 / FQDN / IP.

// Quick health check before starting a research session.
// Uses the Tailscale FQDN — works regardless of MagicDNS resolver state.
const FIRECRAWL_BASE = "http://littleblack.tail0f299b.ts.net:3002";
const res = await fetch(`${FIRECRAWL_BASE}/`);
if (!res.ok) {
  throw new Error(
    `Firecrawl unreachable (${res.status}) — see self-hosted-operations.md and self-hosted-troubleshooting.md`,
  );
}
const banner = await res.json();
if (banner.message !== "Firecrawl API") {
  throw new Error(
    `Unexpected root response: ${JSON.stringify(banner).slice(0, 200)}`,
  );
}

For a true end-to-end probe (proves the full search/scrape stack works, not just the HTTP listener), POST /v1/scrape against https://example.com and check success: true:

curl -s --max-time 15 -X POST \
  "http://littleblack.tail0f299b.ts.net:3002/v1/scrape" \
  -H 'Content-Type: application/json' \
  -d '{"url":"https://example.com","formats":["markdown"]}' \
  | python3 -c "import sys, json; d=json.load(sys.stdin); print('OK' if d.get('success') else 'FAIL')"

Section 2 — Academic Paper Routing

Route paper retrieval to the most effective method based on source. Full decision tree in [academic-paper-routing.md](./references/academic-paper-routing.md).

Quick Reference

| Source | Best Method | Fallback | | ----------------- | ------------------------------------- | ------------------------- | | arxiv.org | Direct HTML (/html/ID) | Firecrawl /v1/scrape | | Semantic Scholar | API (api.semanticscholar.org) | Firecrawl search by title | | ACL Anthology | Firecrawl /v1/scrape | Direct PDF download | | NeurIPS/ICML/ICLR | Firecrawl /v1/scrape with waitFor | Search by title | | IEEE Xplore | Firecrawl with waitFor: 3000 | Author's website | | ACM DL | Firecrawl with waitFor: 3000 | Author's website | | Author blogs | Jina Reader (r.jina.ai) | Firecrawl /v1/scrape | | Google Scholar | Firecrawl /v1/search | Direct search query |

DOI Resolution

// DOI → publisher URL → route to appropriate scraper
const res = await fetch(`https://doi.org/${doi}`, { redirect: "follow" });
const publisherUrl = res.url; // e.g., https://dl.acm.org/doi/10.1145/...
// Then route publisherUrl through the decision tree above

Section 3 — Recursive Research Protocol

The iterative search → extract → recurse → synthesize pattern. Full step-by-step protocol in [recursive-research-protocol.md](./references/recursive-research-protocol.md).

Algorithm Overview

deepResearch(topic, breadth=4, depth=2, concurrency=2):
   1. Generate N search queries (N = breadth) from topic + prior learnings
   2. For each query (via p-limit concurrency):
      a. Firecrawl /v1/search → get results
      b. PERSIST each raw result to docs/research/corpus/
      c. Extract learnings + follow-up questions
   3. For each follow-up question:
      → Recurse with breadth=ceil(breadth/2), depth=depth-1
   4. Base case: depth=0 → return accumulated learnings
   5. Synthesize final report from all learnings
   6. Write session report to docs/research/sessions/

Default Parameters (from working implementation)

| Parameter | Default | Max | Rationale | | ------------- | ------- | --- | ------------------------------------------------------- | | breadth | 4 | — | Number of parallel search queries per level | | depth | 2 | 5 | Recursion levels (depth > 5 yields diminishing returns) | | concurrency | 2 | — | Parallel Firecrawl requests (self-hosted, be gentle) | | limit | 5 | — | Results per search query | | timeout | 15000ms | — | Per-search timeout |

Token Budget

Each search returns up to 5 pages. Trim each page to ~25,000 tokens before LLM processing:

function trimToTokenLimit(text: string, maxTokens: number): string {
  if (!text) return "";
  const estimatedTokens = Math.ceil(text.length / 3.5);
  if (estimatedTokens "          # or any JS-rendered page
NAME="chatgpt-metric-stack-2026-05-27"        # slug — NO whitespace or special chars

# URL-encode the target (avoid Python's trailing newline — use end='')
ENC=$(python3 -c "import urllib.parse,sys; print(urllib.parse.quote(sys.argv[1], safe=''), end='')" "$URL

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [terrylica](https://github.com/terrylica)
- **Source:** [terrylica/cc-skills](https://github.com/terrylica/cc-skills)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.