AgentStack
SKILL verified MIT Self-run

Download Gated Pdfs

skill-kennethkhoocy-legal-scholarship-skills-download-gated-pdfs · by kennethkhoocy

|

No reviews yet
0 installs
0 views
view→install

Install

$ agentstack add skill-kennethkhoocy-legal-scholarship-skills-download-gated-pdfs

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Download Gated Pdfs? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Download bot-gated PDFs via Wayback id_

Problem

Many think-tank and publisher sites (taxpolicycenter.org, urban.org, SSRN delivery) serve an HTML bot-challenge page instead of the PDF to non-browser clients. A browser User-Agent header does not help. The downloaded "PDF" is actually HTML.

Context / Trigger Conditions

  • curl -o file.pdf succeeds but the file starts with `id_/" -o out.pdf

`` is any year likely to have a snapshot (e.g. publication year); Wayback redirects to the nearest capture. The id_` suffix after the timestamp is what requests the untouched original.

  1. Verify the download with pypdf — a bot page fails immediately:

``python from pypdf import PdfReader r = PdfReader("out.pdf"); print(len(r.pages), "pages") ``

  1. If Wayback has no capture, fall back to: another mirror found via search

(Exa/Firecrawl), or Firecrawl scrape for the parsed text when the binary is not strictly needed.

Verification

PdfReader opens the file and reports a plausible page count; first-page text matches the expected title.

Example

Verified 2026-07-15: taxpolicycenter.org/sites/default/files/publication/165884/ssrn-id4797771.pdf and urban.org/sites/default/files/publication/80621/2000790-...pdf both bot-gated to direct curl (with UA), both downloaded intact via https://web.archive.org/web/2024id_/ and .../web/2023id_/ (18 and 12 pages).

Notes

  • Government data hosts (e.g. ticdata.treasury.gov) are usually NOT gated — try direct

curl first; Wayback is the fallback, not the default.

  • Wayback captures can be stale for frequently-revised documents; check the snapshot date

if currency matters.

  • See also: the pdf skill (parsing/extraction after download) and

firecrawl:firecrawl-scrape (parsed markdown when the binary is unnecessary).

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.