AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Eceee News Fulltext Fetch

skill-tiangong-ai-agent-skills-eceee-news-fulltext-fetch · by tiangong-ai

Discover article URLs from https://www.eceee.org/all-news/ and extract/persist full article text into SQLite with retry-safe incremental sync. Use when building or maintaining an eceee news fulltext corpus for downstream search, indexing, or summarization.

— No reviews yet
0 installs
0 views
— view→install

Install

$ agentstack add skill-tiangong-ai-agent-skills-eceee-news-fulltext-fetch

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ✓ Network access No
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ✓ Environment & secrets No
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-tiangong-ai-agent-skills-eceee-news-fulltext-fetch)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● 12d ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Eceee News Fulltext Fetch? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

eceee News Fulltext Fetch

Core Goal

  • Discover news article URLs from https://www.eceee.org/all-news/.
  • Persist discovered entry metadata into SQLite.
  • Fetch and extract article body text from each entry page.
  • Persist status and text in a companion table (entry_content) with retry-safe updates.

Triggering Conditions

  • Receive a request to extract full text from eceee news archive pages.
  • Receive a request to run incremental fulltext sync for eceee news links.
  • Need a resilient local SQLite queue for discovery + extraction + retries.

Workflow

  1. Initialize database.
export ECEEE_NEWS_DB_PATH="/absolute/path/to/eceee_news.db"
python3 scripts/fulltext_fetch.py init-db --db "$ECEEE_NEWS_DB_PATH"
  1. Discover links and fetch fulltext incrementally.
python3 scripts/fulltext_fetch.py sync \
  --db "$ECEEE_NEWS_DB_PATH" \
  --index-url "https://www.eceee.org/all-news/" \
  --limit 50 \
  --min-chars 180
  1. Discover only (refresh URL catalog without fetching bodies).
python3 scripts/fulltext_fetch.py sync \
  --db "$ECEEE_NEWS_DB_PATH" \
  --discover-only
  1. Fetch one entry on demand.
python3 scripts/fulltext_fetch.py fetch-entry \
  --db "$ECEEE_NEWS_DB_PATH" \
  --entry-id 123

Or by URL:

python3 scripts/fulltext_fetch.py fetch-entry \
  --db "$ECEEE_NEWS_DB_PATH" \
  --url "https://www.eceee.org/all-news/news/example-slug/"
  1. Inspect stored state.
python3 scripts/fulltext_fetch.py list-entries --db "$ECEEE_NEWS_DB_PATH" --limit 100
python3 scripts/fulltext_fetch.py list-content --db "$ECEEE_NEWS_DB_PATH" --status ready --limit 100

Data Contract

  • entries table stores discovery metadata:
  • url, title, published_at
  • discovered_at, last_seen_at
  • entry_content table stores extraction result (one row per entry_id):
  • source_url, final_url, http_status
  • extractor (trafilatura, html-parser, or none)
  • content_text, content_hash, content_length
  • status (ready or failed)
  • retry fields + timestamps

Extraction and Update Rules

  • Discovery source is https://www.eceee.org/all-news/, extracting anchor tags with class newslink under /all-news/news/.
  • Fulltext extraction uses article main content region (mainContentColumn) and removes related-news/share blocks.
  • Extraction path:
  1. trafilatura (if installed and not disabled)
  2. built-in HTML parser fallback
  • Upsert by entry_id:
  • Success: set ready, write text/hash/length, reset retry counters.
  • Failure with existing ready content: keep old content, update error/retry metadata.
  • Failure without ready content: set failed, increment retries, set next_retry_at.

Configurable Parameters

  • --db
  • ECEEE_NEWS_DB_PATH
  • --index-url
  • --discover-only
  • --limit
  • --force
  • --only-failed
  • --since-date
  • --refetch-days
  • --oldest-first
  • --timeout
  • --max-bytes
  • --min-chars
  • --max-retries
  • --retry-backoff-minutes
  • --user-agent
  • --disable-trafilatura
  • --fail-on-errors

Error Handling

  • Index fetch/parse failure returns actionable error.
  • HTTP/network/content-type failures are recorded per entry and do not stop the whole sync batch.
  • Short extracted text (< --min-chars) is treated as failed to avoid low-quality bodies.
  • Retry queue is controlled via max_retries + exponential backoff.

References

  • references/schema.md
  • references/fetch-rules.md

Assets

  • assets/config.example.json

Scripts

  • scripts/fulltext_fetch.py

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.