AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Pdf

skill-costa-marcello-skillkit-pdf · by costa-marcello

Extracts text and tables from PDFs, creates new documents, merges/splits files, fills forms, and converts markdown to PDF. Use when working with PDF files, when the user mentions PDFs, document extraction, form filling, OCR, or markdown-to-PDF conversion.

No reviews yet
0 installs
21 views
0.0% view→install

Install

$ agentstack add skill-costa-marcello-skillkit-pdf

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-costa-marcello-skillkit-pdf)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Pdf? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

PDF Processing

Step 1: Identify the Task

Match the user's request to one workflow. Default to Extract Text when unclear.

| Task | Primary Tool | Fallback | |------|-------------|----------| | Extract text | pdfplumber | pdftotext CLI | | Extract tables to DataFrame | pdfplumber + pandas | -- | | OCR scanned PDFs | pytesseract + pdf2image | -- | | Create new PDF | reportlab (Platypus for complex, Canvas for simple) | -- | | Merge/split/rotate | pypdf | qpdf CLI | | Add watermark | pypdf merge_page() | -- | | Password protect/decrypt | pypdf encrypt()/qpdf | -- | | Fill PDF forms | Read references/forms.md and follow its steps | -- | | Markdown to PDF | python scripts/md_to_pdf.py input.md output.pdf | -- | | Batch markdown to PDF | python scripts/batch_convert.py *.md --output-dir ./pdfs/ | -- | | Extract images | pdfimages -j input.pdf output_prefix (poppler-utils) | pypdfium2 | | Render pages to images | pypdfium2 | pdftoppm CLI |

Step 2: Install Dependencies

Check what is available before writing code:

pip list 2>/dev/null | grep -iE "pypdf|pdfplumber|reportlab|pypdfium2|weasyprint|pytesseract|pdf2image"
which pdftotext qpdf pdftk pdfimages 2>/dev/null

Install only what the task needs. Do not install everything.

| Task | Install | |------|---------| | Text/table extraction | pip install pdfplumber | | OCR | pip install pytesseract pdf2image + system poppler | | Merge/split/rotate/encrypt | pip install pypdf | | Create PDFs | pip install reportlab | | Render to images | pip install pypdfium2 | | Markdown to PDF | pip install weasyprint markdown |

Step 3: Execute the Workflow

Use the code patterns in references/cookbook.md for implementation details. Use references/advanced-features.md for pypdfium2, pdf-lib (JS), and advanced CLI operations.

Validation Checkpoints

Run these checks at each stage:

After extraction:

text = page.extract_text()
if not text or len(text.strip()) 

**User:** "Extract all the tables from this PDF and save as Excel"

**Steps:**
1. Install pdfplumber and pandas.
2. Extract tables from each page.
3. Validate each table has rows before saving.
4. Combine and export to Excel.

```python
import pdfplumber
import pandas as pd

with pdfplumber.open("document.pdf") as pdf:
    all_tables = []
    for i, page in enumerate(pdf.pages):
        tables = page.extract_tables()
        for table in tables:
            if table and len(table) > 1:
                df = pd.DataFrame(table[1:], columns=table[0])
                all_tables.append(df)
                print(f"Page {i+1}: extracted table with {len(df)} rows")
    if all_tables:
        combined = pd.concat(all_tables, ignore_index=True)
        combined.to_excel("tables.xlsx", index=False)
        print(f"Saved {len(combined)} total rows to tables.xlsx")
    else:
        print("No tables found in this PDF")

User: "Merge these three PDFs into one"

from pypdf import PdfWriter, PdfReader

writer = PdfWriter()
for pdf_file in ["doc1.pdf", "doc2.pdf", "doc3.pdf"]:
    reader = PdfReader(pdf_file)
    for page in reader.pages:
        writer.add_page(page)
    print(f"Added {len(reader.pages)} pages from {pdf_file}")

with open("merged.pdf", "wb") as output:
    writer.write(output)

# Verify
result = PdfReader("merged.pdf")
print(f"Merged PDF has {len(result.pages)} pages")

User: "This PDF is scanned, I need the text"

import pytesseract
from pdf2image import convert_from_path

images = convert_from_path("scanned.pdf", dpi=300)
text = ""
for i, image in enumerate(images):
    page_text = pytesseract.image_to_string(image)
    text += page_text
    print(f"Page {i+1}: extracted {len(page_text)} characters")

with open("extracted.txt", "w") as f:
    f.write(text)
print(f"Total: {len(text)} characters. Review for OCR errors.")

User: "Convert this markdown file to PDF"

# macOS may need: export DYLD_LIBRARY_PATH=$(brew --prefix)/lib
python scripts/md_to_pdf.py report.md report.pdf

Features: A4 pages, proper margins, table/code support, Chinese font fallback. For batch conversion: python scripts/batch_convert.py *.md --output-dir ./pdfs/

User: "Fill out this PDF form with my details"

Follow the complete workflow in references/forms.md. Summary:

  1. Check if the PDF has fillable fields: python scripts/check_fillable_fields.py form.pdf
  2. If fillable: extract field info, create values JSON, run fill script.
  3. If not fillable: convert to images, identify fields visually, create bounding boxes, validate, fill with annotations.

References

| File | Purpose | |------|---------| | references/cookbook.md | Python and CLI code patterns for all PDF operations | | references/advanced-features.md | pypdfium2, pdf-lib (JS), advanced CLI, performance tips | | references/forms.md | Complete form-filling workflow (fillable and non-fillable PDFs) |

Scripts

| Script | Purpose | Usage | |--------|---------|-------| | scripts/md_to_pdf.py | Markdown to PDF with Chinese font support | python scripts/md_to_pdf.py input.md [output.pdf] | | scripts/batch_convert.py | Batch markdown to PDF | python scripts/batch_convert.py *.md [--output-dir dir] | | scripts/check_fillable_fields.py | Check if PDF has fillable form fields | python scripts/check_fillable_fields.py input.pdf | | scripts/extract_form_field_info.py | Extract form field metadata to JSON | python scripts/extract_form_field_info.py input.pdf output.json | | scripts/fill_fillable_fields.py | Fill fillable PDF form fields | python scripts/fill_fillable_fields.py input.pdf values.json output.pdf | | scripts/fill_pdf_form_with_annotations.py | Fill non-fillable PDFs with text annotations | python scripts/fill_pdf_form_with_annotations.py input.pdf fields.json output.pdf | | scripts/convert_pdf_to_images.py | Convert PDF pages to PNG images | python scripts/convert_pdf_to_images.py input.pdf output_dir | | scripts/create_validation_image.py | Create bounding box validation images | python scripts/create_validation_image.py page_num fields.json input.png output.png | | scripts/check_bounding_boxes.py | Validate bounding boxes do not overlap | python scripts/check_bounding_boxes.py fields.json |

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.