# Document Skills Pdf

> Extract and process PDF documents for data pipelines. Use when parsing research papers, extracting tables from reports, or converting PDFs to structured data.

- **Type:** Skill
- **Install:** `agentstack add skill-ihatesea69-kiro-kit-pdf`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [ihatesea69](https://agentstack.voostack.com/s/ihatesea69)
- **Installs:** 0
- **Category:** [Content & Media](https://agentstack.voostack.com/c/content-and-media)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [ihatesea69](https://github.com/ihatesea69)
- **Source:** https://github.com/ihatesea69/kiro-kit/tree/main/.kiro/skills/document-skills/pdf
- **Website:** https://www.npmjs.com/package/kiro-kit

## Install

```sh
agentstack add skill-ihatesea69-kiro-kit-pdf
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Document Skills - PDF

Activate this skill when working with PDF documents in data/AI workflows.

## When to Use

- Extracting tables from financial reports
- Parsing research papers for literature review
- Converting scanned PDFs to text (OCR)
- Extracting metadata from document collections
- Building document processing pipelines

## Libraries

- **pypdf**: Read/write PDF, extract text and metadata
- **pdfplumber**: Table extraction with spatial awareness
- **PyMuPDF (fitz)**: Fast rendering and text extraction
- **camelot-py**: Table extraction from PDFs
- **pytesseract**: OCR for scanned documents

## Usage

```python
import pdfplumber

with pdfplumber.open("report.pdf") as pdf:
    for page in pdf.pages:
        tables = page.extract_tables()
        text = page.extract_text()

# OCR for scanned documents
import pytesseract
from pdf2image import convert_from_path

images = convert_from_path("scanned.pdf")
text = pytesseract.image_to_string(images[0])
```

## Rules

- Check if PDF is text-based or scanned before processing
- Use pdfplumber for table extraction over regex parsing
- Handle multi-column layouts carefully
- Validate extracted numbers against visual inspection
- Process large PDFs page-by-page to manage memory

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [ihatesea69](https://github.com/ihatesea69)
- **Source:** [ihatesea69/kiro-kit](https://github.com/ihatesea69/kiro-kit)
- **License:** MIT
- **Homepage:** https://www.npmjs.com/package/kiro-kit

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-ihatesea69-kiro-kit-pdf
- Seller: https://agentstack.voostack.com/s/ihatesea69
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
