Install
$ agentstack add skill-lisniuse-taoguba-crawler-skill-taoguba-crawler-skill ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Taoguba Crawler
This skill runs the Taoguba (tgb.cn) article crawlers located in the project root.
Prerequisites
- Python 3 with
requests,beautifulsoup4,python-dotenvinstalled - A
.envfile in the project root containingCOOKIEand optionallyUSER_AGENT
Available Crawlers
1. BBS Crawler (crawler_bbs.py)
Crawl the forum board at tgb.cn/bbs/1/1 using HTML scraping.
python crawler_bbs.py
- Extracts article list by parsing
a.overhide.mw300elements - Gets each article's main post and author replies
- Downloads images and embeds them as base64 in HTML
- Outputs:
output/bbs_YYYY-MM-DD.jsonandoutput/bbs_YYYY-MM-DD_HHMMSS.html
2. Home Crawler (crawler_home.py)
Crawl the homepage recommendations via JSON API (/newIndex/getZh).
python crawler_home.py
- Fetches articles from the JSON API (default 2 pages)
- Same content extraction and HTML generation as BBS crawler
- Outputs:
output/home_YYYY-MM-DD.jsonandoutput/home_YYYY-MM-DD_HHMMSS.html
Common Workflow
To run both crawlers:
python crawler_bbs.py && python crawler_home.py
Key Implementation Details
- Authentication: Both scripts read
COOKIEfrom.envviapython-dotenv - Rate limiting: 0.5-1s delay between requests to avoid being blocked
- Image handling: Images are downloaded and embedded as base64 in the HTML output
- Article content: Extracts main post (
#first) and author replies (.comment-datawith author badge) - Output directory: All results saved to
output/folder
Scripts
The crawler scripts are bundled in scripts/:
scripts/crawler_bbs.py- BBS forum crawler (HTML scraping)scripts/crawler_home.py- Homepage crawler (JSON API)
To run the bundled scripts directly:
python scripts/crawler_bbs.py
python scripts/crawler_home.py
Troubleshooting
- If no articles are returned, check that
.envcontains a validCOOKIEvalue - If image downloads fail, the HTML will show error messages inline
- Network timeouts default to 10-15 seconds per request
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: lisniuse
- Source: lisniuse/taoguba-crawler-skill
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.