# Pdf Reading

> Extract text, tables, and structured information from PDF documents using pdfplumber, PyPDF2, or pdftotext command-line tools.

- **Type:** Skill
- **Install:** `agentstack add skill-xushuwenn-redact-pdf-reading`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [XuShuwenn](https://agentstack.voostack.com/s/xushuwenn)
- **Installs:** 0
- **Category:** [Content & Media](https://agentstack.voostack.com/c/content-and-media)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [XuShuwenn](https://github.com/XuShuwenn)
- **Source:** https://github.com/XuShuwenn/RedAct/tree/main/captracebench/libs/artifact-runner/tasks/nodemedic-demo/environment/skills/pdf-reading

## Install

```sh
agentstack add skill-xushuwenn-redact-pdf-reading
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# PDF Reading Skill

This skill helps agents extract information from PDF documents.

## Tools Available

The environment has these PDF tools pre-installed:
- `pdfplumber` - Python library for precise text/table extraction
- `PyPDF2` - Python library for PDF manipulation
- `pdftotext` - Command-line tool from poppler-utils

## Quick Extraction

### Command Line (Fast)

```bash
# Extract all text from a PDF
pdftotext /root/artifacts/paper.pdf -

# Extract specific pages
pdftotext -f 1 -l 3 /root/artifacts/paper.pdf -

# Extract to a file
pdftotext /root/artifacts/paper.pdf /tmp/paper.txt
```

### Python (More Control)

```python
import pdfplumber
from pathlib import Path

def extract_pdf_text(pdf_path: str) -> str:
    """Extract all text from a PDF file."""
    text_parts = []
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            text = page.extract_text()
            if text:
                text_parts.append(text)
    return "\n\n".join(text_parts)

# Usage
text = extract_pdf_text("/root/artifacts/paper.pdf")
print(text)
```

## Extracting Specific Information

### Find Commands in Text

```python
import re

def find_commands(text: str) -> list:
    """Extract shell commands from text."""
    # Look for common command patterns
    patterns = [
        r'docker run[^\n]+',
        r'\$[^\n]+',
        r'--package=[^\s]+\s+--version=[^\s]+',
    ]
    commands = []
    for pattern in patterns:
        commands.extend(re.findall(pattern, text))
    return commands

text = extract_pdf_text("/root/artifacts/paper.pdf")
commands = find_commands(text)
```

### Extract Tables

```python
import pdfplumber

def extract_tables(pdf_path: str) -> list:
    """Extract all tables from a PDF."""
    tables = []
    with pdfplumber.open(pdf_path) as pdf:
        for i, page in enumerate(pdf.pages):
            page_tables = page.extract_tables()
            for table in page_tables:
                tables.append({
                    "page": i + 1,
                    "data": table
                })
    return tables
```

### Find Package Information

```python
import re

def find_package_info(text: str) -> list:
    """Find npm package references (name@version)."""
    # Match patterns like node-rsync@1.0.3
    pattern = r'([a-z0-9-]+)@(\d+\.\d+\.\d+)'
    matches = re.findall(pattern, text.lower())
    return [{"name": m[0], "version": m[1]} for m in matches]
```

## Tips

1. **Check available PDFs first**:
   ```bash
   ls -la /root/artifacts/
   ```

2. **Preview before full extraction**:
   ```bash
   pdftotext /root/artifacts/paper.pdf - | head -100
   ```

3. **Handle multi-column layouts**: pdfplumber handles them better than pdftotext

4. **For structured data**: Look for JSON blocks in the text:
   ```python
   import json
   import re

   json_blocks = re.findall(r'\{[^{}]*\}', text)
   for block in json_blocks:
       try:
           data = json.loads(block)
           print(data)
       except json.JSONDecodeError:
           pass
   ```

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [XuShuwenn](https://github.com/XuShuwenn)
- **Source:** [XuShuwenn/RedAct](https://github.com/XuShuwenn/RedAct)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** yes
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-xushuwenn-redact-pdf-reading
- Seller: https://agentstack.voostack.com/s/xushuwenn
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
