# Collecting Sources

> Collects research sources with quality evaluation. Use when gathering or finding sources, building a source library, searching for academic papers, or performing RADAR assessment.

- **Type:** Skill
- **Install:** `agentstack add skill-isvlasov-rageatc-oss-collecting-sources`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [isvlasov](https://agentstack.voostack.com/s/isvlasov)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [isvlasov](https://github.com/isvlasov)
- **Source:** https://github.com/isvlasov/rageatc-oss/tree/main/plugins/rageatc-core-oss/skills/collecting-sources

## Install

```sh
agentstack add skill-isvlasov-rageatc-oss-collecting-sources
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Collecting Sources

## Purpose

Enable systematic, high-quality source collection for research workflows by discovering relevant sources using domain-appropriate tools, evaluating quality using RADAR framework, storing sources locally with complete metadata, and creating human-readable source indices.

This skill supports the first phase of two-phase research workflows:
- **Phase 1 (this skill)**: Discover, evaluate, and store sources with metadata
- **Phase 2**: Synthesise findings from collected sources
- **Phase 3 (optional)**: Fact-check claims in synthesis

## When to Use This Skill

Use this skill when:
- Beginning research that requires collecting and storing sources with metadata
- User requests "collect sources", "gather sources", "find sources for research"
- Starting systematic research on unfamiliar topics requiring source quality assessment
- Building a source library with provenance tracking for reproducibility
- Conducting academic research requiring structured source management
- Working within two-phase research workflow (collection → synthesis)

**Do NOT use this skill when:**
- Conducting simple web searches for immediate answers (use WebSearch directly)
- Research requires only Claude's training knowledge
- User wants synthesis without separate collection phase
- Quick fact lookup rather than systematic research

## Required Inputs

Before beginning source collection, ensure these inputs are provided:

**Always required:**
- [ ] **Research question(s)** - Specific questions or topics to investigate
- [ ] **Task ID** - Workspace identifier (for `work//` directory structure)
- [ ] **Minimum source count** - Target number of sources (default: 8-15)

**Context-dependent:**
- [ ] **Domain context** - Academic vs general; if academic: biomedical/CS/physics/social sciences/general (for tool selection)
- [ ] **Quality threshold** - Minimum acceptable reliability score (default: 0.6)
- [ ] **Source type preferences** - Prioritise academic papers, blogs, documentation, etc.
- [ ] **Time constraints** - Affects search breadth and depth

If inputs are missing, ask the orchestrator before proceeding.

## Outputs Produced

This skill produces:

1. **Sources directory structure**:
   ```
   work//sources/
   ├── papers/          # Academic papers (PDF or plain text)
   ├── web/             # Web pages (Markdown)
   ├── blogs/           # Blog posts (Markdown)
   ├── docs/            # Documentation (Markdown/HTML)
   └── source_index.md  # Human-readable catalogue
   ```

2. **Source files**: Content stored in appropriate formats (PDF/plain text for papers, Markdown for web pages)

3. **Metadata files**: `.meta.yaml` companion files with complete metadata conforming to schema v1.0

4. **Source index**: `source_index.md` cataloguing all sources with statistics and grouped listings

5. **Collection summary**: Report with statistics, limitations, and handoff notes

## Core Workflow

### Phase 0: Setup and Input Validation

**Goal**: Prepare workspace and validate inputs

1. **Validate required inputs**:
   - Confirm research question(s) received
   - Confirm task_id received
   - Check for domain context hints in research question

2. **Create workspace structure**:
   ```bash
   mkdir -p work//sources/{papers,web,blogs,docs}
   ```

3. **Initialise tracking**:
   - Set source counter to 1 (for src_001, src_002, etc.)
   - Create collection log for provenance tracking
   - Check if `source_index.md` exists from previous collection (for duplicate detection)

**Checkpoint**: Workspace exists, inputs validated, ready to begin discovery.

---

### Phase 1: Domain Detection and Query Formulation

**Goal**: Identify research domain and formulate effective search queries

#### Domain Detection

Apply domain detection heuristics to research question:

**Academic domain indicators:**
- Keywords: "paper", "study", "research", "peer-reviewed", "journal", "conference"
- Biomedical: "medicine", "clinical", "disease", "drug", "genomic", "patient", "healthcare"
- CS/Physics: "algorithm", "neural network", "machine learning", "quantum", "physics", "computer science"
- Social sciences: "policy", "governance", "sociology", "political science", "economics", "education"
- Academic URLs provided: arxiv.org, doi.org, pubmed.gov, semanticscholar.org

**General web indicators:**
- Keywords: "tutorial", "guide", "how-to", "blog", "documentation", "best practices"
- No academic terminology
- Practical/applied focus

#### Query Formulation Strategy

Based on domain detection:

**For Academic Research:**

Academic paper retrieval uses a **two-step workflow**:
1. **Discovery** — Find papers via metadata APIs or web search
2. **Retrieval** — Fetch full text using multi-source strategy (detailed in Phase 3)

**Discovery queries:**

1. **Biomedical topics** → Site-specific searches + CORE API awareness
   - Example: `site:pubmed.ncbi.nlm.nih.gov clinical trial metadata`
   - Example: `clinical trial metadata standards peer-reviewed`
   - Note DOIs/titles for full-text retrieval via CORE or PubMed Central

2. **CS/Physics topics** → Site-specific searches + arXiv awareness
   - Example: `site:arxiv.org transformer architecture`
   - Example: `site:scholar.google.com neural network efficiency`
   - Note arXiv IDs for direct PDF retrieval

3. **Social sciences** → Metadata APIs + CORE fallback
   - Example: `governance NGO policy research Egypt`
   - Example: `site:scholar.google.com civil society regulation`
   - Note: Social sciences have lower OA coverage (~33% vs ~66% STEM); expect higher metadata-only rate

4. **General academic** → Multiple search strategies
   - WebSearch with academic keywords: `research paper [topic]`
   - Site-specific: `site:scholar.google.com [topic]`
   - Extract DOIs/titles for full-text retrieval

**For General Web Research:**
- Use WebSearch with standard queries
- Target authoritative sources: official docs, expert blogs, reputable sites
- Use site-specific searches when known: `site:docs.python.org async`

**Checkpoint**: Domain identified, search queries formulated, ready for discovery.

---

### Phase 2: Source Discovery

**Goal**: Discover relevant candidate sources and extract identifiers for full-text retrieval

#### Discovery Strategy

**For academic research:**

The goal is to find papers and extract **identifiers** (DOI, arXiv ID, PubMed ID, title) for full-text retrieval, not to fetch full text during discovery.

1. **Formulate metadata queries**:
   - Extract key terms from research question
   - Use domain-specific terminology
   - Consider alternative phrasings
   - Use site-specific searches where applicable

2. **Execute discovery searches**:
   - Run 3-5 targeted WebSearch queries
   - Use domain-specific site searches (site:arxiv.org, site:pubmed.gov, site:scholar.google.com)
   - Focus on finding paper metadata: titles, authors, DOIs, abstracts
   - Collect **identifiers** for full-text retrieval:
     - DOI (for CORE API, Unpaywall MCP)
     - arXiv ID (for arXiv direct retrieval)
     - PubMed ID (for PubMed Central)
     - Full paper title (fallback for CORE API search)

3. **Collect candidate metadata**:
   - Paper titles and DOIs
   - Authors and publication dates
   - Venue (journal/conference)
   - Abstract excerpts
   - Repository identifiers (arXiv ID, PMID, etc.)

**For general web research:**

1. **Formulate search queries**:
   - Extract key terms from research question
   - Use domain-specific terminology
   - Consider alternative phrasings

2. **Execute searches**:
   - Run 3-5 targeted WebSearch queries
   - Use site-specific searches where applicable
   - Collect URLs and metadata from search results
   - Aim for diverse source types

3. **Collect candidate metadata**:
   - URLs and titles
   - Authors and publication dates (where available)
   - Site names and descriptions

**Target**: 15-25 candidate sources from diverse source types (will filter to 8-15 based on quality)

**Checkpoint**: Candidate sources identified with identifiers (DOI/arXiv ID/title) for academic papers, URLs for web sources.

---

### Phase 3: Source Retrieval and Storage

**Goal**: Fetch content using multi-source strategy and store in appropriate formats

For each candidate source:

#### 3.1: Check for Duplicates

**Before fetching**, check if source already exists:

1. **Load existing sources** (if `source_index.md` exists from previous collection)
2. **Check URL match**: Compare candidate URL against existing source URLs
3. **Check DOI match**: For academic papers, compare candidate DOI against `doi` fields in existing `.meta.yaml` files in the sources directory
4. **If exists**: Skip fetching; log as "already collected"
5. **If differs**: Continue to fetch (content_hash comparison will catch duplicate content later)

This avoids re-fetching sources already collected in previous sessions.

#### 3.2: Fetch Content — Multi-Source Strategy

**For academic papers:**

Use a **structured fallback chain** to maximise full-text retrieval:

**Retrieval Chain:**
1. **CORE API** (largest corpus, plain text in JSON)
2. **Unpaywall MCP** (OA PDF extraction)
3. **Domain-specific repository** (arXiv, PubMed Central)
4. **WebFetch** (publisher sites, repository landing pages)
5. **Metadata-only** (last resort for paywalled content)

**Detailed retrieval strategy:**

**Step 1: CORE API** (primary for all academic papers)

CORE API provides direct access to 46M full-text papers as plain text in JSON. No PDF extraction needed.

**When to use:**
- Always try first for academic papers with DOI or title
- Covers all academic domains (STEM, social sciences, humanities)
- Returns plain text extracted from PDFs during indexing

**How to use:**
1. **Query by DOI**:
   ```
   WebFetch: https://api.core.ac.uk/v3/search/works?q=doi:[DOI]
   ```
   - Response: JSON with `fullText` field containing plain text
   - Extract `fullText` from JSON response

2. **Query by title** (if no DOI):
   ```
   WebFetch: https://api.core.ac.uk/v3/search/works?q=title:"[exact title]"
   ```
   - Review results for title match
   - Extract `fullText` from matching paper

3. **Extract plain text**:
   - If `fullText` field present and non-empty → save as `.txt` file
   - Plain text is pre-extracted via Apache PDFBox; ready to use
   - If `fullText` empty or absent → proceed to Step 2 (Unpaywall)

**Rate limits**: 1 batch request or 5 single requests per 10 seconds (free tier)

**Step 2: Unpaywall MCP** (fallback for OA papers)

If CORE API fails or returns no full text, use Unpaywall MCP tools to find and extract OA PDFs.

**MCP tools available:**
- `unpaywall_search_titles` — Search by title, returns OA status
- `unpaywall_get_fulltext_links` — Get best OA PDF URL by DOI
- `unpaywall_fetch_pdf_text` — Download PDF and extract text (configurable truncation)

**How to use:**
1. **Get OA PDF URL** (if DOI available):
   ```
   unpaywall_get_fulltext_links(doi="10.1234/example")
   ```
   - Returns `best_oa_location.url_for_pdf` if available
   - Proceed to extraction if PDF URL found

2. **Search by title** (if no DOI):
   ```
   unpaywall_search_titles(query="exact paper title")
   ```
   - Review results for title match
   - Extract DOI if found, then use `unpaywall_get_fulltext_links`

3. **Extract PDF text**:
   ```
   unpaywall_fetch_pdf_text(doi="10.1234/example", truncate_chars=50000)
   ```
   - Downloads PDF and extracts text
   - Configure `truncate_chars` to 50,000 for comprehensive coverage
   - Minimum 1,000 characters enforced
   - Average papers: 20,000-30,000 chars; surveys: 90,000-100,000 chars

4. **Save extracted text**:
   - If extraction succeeds → save as `.txt` file
   - If extraction fails (paywalled, no OA version) → proceed to Step 3

**Rate limits**: Respect Unpaywall's 100,000 calls/day limit

**Step 3: Domain-Specific Repository** (for known domains)

If CORE and Unpaywall fail, try domain-specific repositories with near-complete coverage.

**3a. arXiv** (for CS, physics, mathematics, quantitative biology):

**When to use:**
- arXiv ID detected in metadata
- CS/physics/maths keywords in research question
- Paper published in arXiv venue

**How to use:**
1. **Prefer arXiv MCP server** (if available) for metadata + text extraction
2. **Alternative — try CORE API first**: arXiv papers are often in CORE's corpus. Query CORE by arXiv ID or title to get plain text without PDF extraction
3. **Last resort — direct PDF download**:
   ```
   WebFetch: https://export.arxiv.org/pdf/[arxiv_id].pdf
   ```
   - **Note:** WebFetch cannot extract text from PDFs. This downloads the binary PDF for storage only
   - Suggested rate: 4 requests/second with 1-second sleep per burst

4. **Store PDF**:
   - Save as `.pdf` in `sources/papers/`
   - Note in metadata: `Retrieved from arXiv`

**3b. PubMed Central** (for biomedical, life sciences):

**When to use:**
- PubMed ID detected in metadata
- Biomedical keywords in research question
- Paper published in biomedical journal

**How to use:**
1. **WebFetch landing page**:
   ```
   WebFetch: https://pmc.ncbi.nlm.nih.gov/articles/PMC[PMCID]/
   ```
   - Extracts full-text HTML
   - PMC provides full articles, not just abstracts

2. **Store as HTML or Markdown**:
   - Save HTML or convert to Markdown
   - Note in metadata: `Retrieved from PubMed Central`

**Step 4: WebFetch** (general fallback for open-access content)

If all targeted strategies fail, try standard web retrieval.

**How to use:**
1. **Use WebFetch** on paper landing page URL
2. **Extract what's available**:
   - HTML metadata pages → abstract, bibliographic info
   - Open-access PDFs → download if link found
   - Repository pages → may include full text as HTML

3. **Result**:
   - If full text retrieved → store normally
   - If landing page only → record metadata, proceed to Step 5

**Step 5: Metadata-Only** (last resort for paywalled content)

If all retrieval strategies fail (paywalled, not in OA repositories):

1. **Record metadata** from abstract/landing page:
   - Title, authors, DOI, abstract, venue, publication date
   - Extract from WebFetch of landing page

2. **Attempt open-access alternatives**:
   - Search for preprint: `site:arxiv.org [title]`
   - Search for preprint: `site:biorxiv.org [title]`
   - Check author websites: `[author name] [title] pdf`

3. **If unavailable**:
   - Create metadata file with all available fields
   - Set `file_path: "unavailable"`
   - Set `content_hash: "unavailable"`
   - Note in `provenance.notes`: "Paywalled; full text unavailable via CORE, Unpaywall, domain repositories; metadata only"
   - Flag in source index: "Metadata only (paywalled)"

**Social sciences caveat:**

Social sciences have substantially lower OA coverage (~33%) compared to STEM (~66%). For social science research:
- Expect higher metadata-only rates (50-60% paywalled not uncommon)
- SSRN (1.74M social sciences preprints) lacks API access; cannot retrieve programmatically
- No domain-specific repository equivalent to arXiv/PubMed Central for social sciences
- CORE + Unpaywall are primary strategies; calibrate expectations accordingly

**For web pages/blogs/documentation:**

Use WebFetch to retrieve content:
- WebFetch converts HTML to Markdown automatically
- Store as Markdown (`.md`)
- Preserve HTML only if formatting is critical

#### 3.3: Determine Source Type

Classify source using URL and content patterns:

- `source_type: "academic_paper"` - Peer-reviewed papers, preprints, conference papers
- `source_type: "web_page"` - General web content, articles, guides
- `source_type: "blog"` - Blog posts, Substack, Medium, personal sites
- `source_type: "documentation"` - Official project docs, API references, specifications

#### 3.4: Select Storage Format

Match format to source type and retrieval method:

| Source Type | Primary Format | Use Case |
|-------------|---------------|----------|
| Academic paper (CORE API) | Plain text (`.txt`) | Pre-extracted text from CORE, ready to use |
| Academic p

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [isvlasov](https://github.com/isvlasov)
- **Source:** [isvlasov/rageatc-oss](https://github.com/isvlasov/rageatc-oss)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-isvlasov-rageatc-oss-collecting-sources
- Seller: https://agentstack.voostack.com/s/isvlasov
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
