Install
$ agentstack add skill-chuongdlb-agent-skills-book-to-knowledge-base ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ● Filesystem access Used
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Book to Knowledge Base — Lossless PDF-to-Markdown Conversion
Purpose
Convert a technical book (PDF chapters) into a structured markdown knowledge base that preserves all mathematical and technical content. The output is a self-contained directory of markdown files — one per chapter, plus an overview and cross-reference index — suitable for AI-assisted retrieval via Read/Grep.
When to Use
Invoke this skill when:
- You have a technical/mathematical book as one or more PDF files
- You want a lossless markdown knowledge base (not a summary or curated extraction)
- The goal is a persistent reference that can be queried by AI agents
- You need chapter-dependency tracking for structured navigation
Not for: Quick summaries, topic-oriented reorganization, or non-technical books. For curated/topic-oriented KB construction, see the book-reader skill.
Key Design Decisions
Chapter-by-chapter, not topic-oriented. Each output file maps 1:1 to a book chapter. This preserves the author's pedagogical structure, makes verification trivial (compare file to PDF), and avoids lossy reorganization. Cross-cutting queries are handled by the cross-reference index.
Lossless, not curated. Include everything: every definition, theorem (with ALL conditions), algorithm, equation, proof strategy, example, and remark. Judgment calls about what to exclude introduce errors. Let the reader/agent decide what's relevant at query time.
Small enough for direct access. Target total KB size under 50K tokens (~8,000 lines). This fits comfortably in Read/Grep access without needing a vector DB or MCP server.
Output Specification
Directory Structure
knowledge_base/
00-overview.md # Reading order, dependency diagram, chapter summary table
01-.md # Chapter 1
02-.md # Chapter 2
...
NN-.md # Last chapter
(NN+1)-appendix.md # Appendix (if book has one)
(NN+2)-cross-reference-index.md # Master index across all chapters
File Naming
- Two-digit zero-padded prefix:
01-,02-, ...,12- - Slug derived from chapter title: lowercase, hyphens for spaces
- Examples:
01-basic-concepts.md,07-temporal-difference-methods.md,11-appendix.md
Chapter File Format
Every chapter file follows this exact structure:
---
chapter:
title:
key_topics: [topic1, topic2, topic3, ...]
depends_on: [1, 2, 3]
required_by: [5, 6, 7]
---
# Chapter N:
> Source: ** (, ), Chapter N, pp. X-Y
> Supplemented by:
> Errata:
## Purpose and Context
---
## N.1
---
## N.2
...
---
## Summary and Key Takeaways
YAML Frontmatter Fields
| Field | Type | Description | |-------|------|-------------| | chapter | int or string | Chapter number, or "appendix", or "overview" | | title | string | Chapter title from the book | | key_topics | string[] | All major topics covered, for search/filtering | | depends_on | int[] | Chapter numbers that are prerequisites | | required_by | int[] | Chapter numbers that depend on this one |
Overview File Format (00-overview.md)
---
type: overview
title: Knowledge Base Overview
---
# Knowledge Base Overview
---
## Reading Order and Dependency Diagram
---
## Chapter Summary Table
| Chapter | Title | Key Topics | Pages | Algorithms Covered |
|---------|-------|------------|-------|--------------------|
| 1 | ... | ... | 1-13 | (none) |
| 2 | ... | ... | 15-34 | ... |
---
## Concept Flow Diagram
---
## File Listing
| File | Description |
|------|-------------|
| 00-overview.md | This file. |
| 01-... | Chapter 1: ... |
---
## Navigation Tips for AI Agents
Cross-Reference Index Format
---
type: synthesis
title: Cross-Reference Index and Algorithm Comparison
---
# Cross-Reference Index and Algorithm Comparison
## 1. Master Concept Index
| Concept | Defined In | Also Used In |
...
## 2. Algorithm Comparison Table
| Algorithm | Type | Model Required? | ... |
...
## 3. Theorem and Lemma Index
| Theorem | Name | One-Line Statement | Chapter | Key Conditions |
...
## 4. Equation Quick-Reference
## 5. Key Transitions
Pipeline
Phase 0: Environment & Inventory
Step 0.1 — Verify PDF reading capability
# Option A: Claude Code Read tool (works if poppler-utils installed)
# Try reading a small PDF first
# Option B: pymupdf fallback
uv run --with pymupdf python3 -c "
import fitz
doc = fitz.open('PATH_TO_PDF')
for i in range(min(3, len(doc))):
print(f'=== PAGE {i+1} ===')
print(doc[i].get_text())
"
CRITICAL: Verify PDF reading produces clean text on 2-3 pages BEFORE proceeding. If math symbols are garbled, try a different extraction method. Do not proceed with broken extraction.
Step 0.2 — Inventory all files
List all PDFs and identify: chapter PDFs, combined PDF, table of contents, appendices, errata, supplementary material (slides, code).
Step 0.3 — Create output directory
mkdir -p /path/to/repo/knowledge_base
Phase 1: Table of Contents & Structure
Read the table of contents, preface, and overview to extract:
- Chapter list with titles and page ranges
- Chapter dependencies (which chapters build on which)
- Book structure (parts, sections per chapter)
- Supplementary material mapping (which slides/code maps to which chapter)
- Errata (corrections to apply during extraction)
Save as a working document. This informs everything that follows.
Phase 2: Chapter-by-Chapter Extraction
Process each chapter using the extraction protocol below. For books with 8+ chapters, use the parallelization strategy (see below).
Extraction Protocol (per chapter)
Read the entire chapter (20 pages at a time for large chapters). Extract ALL of the following:
| Category | What to Extract | Formatting | |----------|----------------|------------| | Definitions | Every formal definition with its number and exact statement | Block with "Definition N.M:" prefix | | Theorems/Lemmas | Complete statements with ALL conditions and conclusions | Block with "Theorem N.M:" prefix; list conditions explicitly | | Proof strategies | Key technique and steps (not line-by-line) | Under theorem, as "Proof sketch:" | | Algorithms | Full pseudocode with input/output/steps | As "Algorithm N.M: Name" with numbered steps | | Equations | Every numbered equation, in LaTeX | $$...\tag{N.M}$$ format | | Examples | All worked examples that illustrate concepts | Under the relevant section | | Remarks | Author's insights, practical advice, common mistakes | As "Remark:" or "Important note:" | | Relationships | How this chapter connects to other chapters | In "Purpose and Context" section and inline forward/back references |
Formatting Rules
- LaTeX notation for all math:
$inline$and$$display$$ - Section numbering mirrors the book:
## N.1,## N.2, etc. - Equation tags: Use
\tag{N.M}matching the book's equation numbers - Tables for structured data (state transition tables, reward tables, comparison tables)
- Horizontal rules (
---) between major sections - Block quotes for source citations:
> Source: *Book Title* ... - Bold for definitions, theorem names, algorithm names
- Forward/back references as inline text: "See Chapter N" or "as shown in equation (M.K)"
- Errata corrections: Apply silently, note in the source citation block
Per-Chapter Quality Gate
Before moving to the next chapter, verify:
- [ ] Every section from the book has a corresponding
## N.Mheading - [ ] All theorems have complete conditions (no missing assumptions)
- [ ] All algorithms have complete pseudocode
- [ ] Equation numbering matches the book
- [ ] YAML frontmatter
depends_onandrequired_byare accurate
Phase 3: Synthesis Files
After all chapters are extracted:
Step 3.1 — Overview file (00-overview.md)
Build from the Phase 1 structure document. Include:
- ASCII dependency diagram (drawn from the
depends_on/required_bymetadata) - Chapter summary table (one row per chapter with: title, key topics, page range, algorithms)
- Concept flow diagram showing the book's intellectual progression
- File listing with one-line descriptions
- Navigation tips for AI agents
Step 3.2 — Cross-reference index
Scan ALL chapter files to build:
- Master concept index: Every major concept → where defined, where used
- Algorithm comparison table: All algorithms in one table with type, model requirement, convergence guarantee, on/off-policy
- Theorem/lemma index: All theorems with one-line statements, chapter, key conditions
- Equation quick-reference: The ~20 most important equations grouped by topic
- Key transitions: Major conceptual shifts across the book (e.g., model-based → model-free, tabular → function approximation)
Phase 4: Verification
Step 4.1 — Completeness audit
For each chapter, compare the markdown file against the PDF:
- Count of definitions matches
- Count of theorems/lemmas matches
- Count of algorithms matches
- All numbered equations present
- No sections skipped
Step 4.2 — Cross-reference integrity
- Every concept in the cross-reference index points to a real section
- Every
depends_on/required_byreference is reciprocal (if Ch 3 depends on Ch 2, then Ch 2'srequired_byincludes 3) - File listing in overview matches actual files
Step 4.3 — Sample query test
Pick 3 known results from the book and verify they can be found via:
- Concept lookup in
cross-reference-index.md - Direct
Grepacrossknowledge_base/*.md - Reading the specific chapter file
Phase 5: Integration
Step 5.1 — Update project CLAUDE.md
Add a "Knowledge Base" section documenting:
- File map (table of files → content)
- How to query the KB (concept lookup, keyword search, chapter dependencies)
- Design decision: no MCP server (KB is small enough for direct Read/Grep)
Step 5.2 — Report
Output a summary:
- Number of files created with line counts
- Total line count and estimated token count
- Any chapters or sections with extraction issues
- Sample query demonstrating the KB works
Parallelization Strategy
For books with 8+ chapters
Read chapters in parallel using background agents. Group by dependency tiers:
Tier 1 (no dependencies): Ch 1, Appendix
Tier 2 (depends on Tier 1): Ch 2, Ch 3
Tier 3 (depends on Tier 2): Ch 4, Ch 5, Ch 6
Tier 4 (depends on Tier 3): Ch 7, Ch 8, Ch 9, Ch 10
Launch 3-5 agents per tier. Each agent receives:
- The chapter PDF (or page range of the combined PDF)
- The extraction protocol (copy the full protocol from this skill)
- The output path
- The errata relevant to this chapter
Agent prompt template:
Read [PDF_PATH] pages [START]-[END] and convert Chapter [N]: "[TITLE]" to markdown.
Follow this exact format:
[Paste the Chapter File Format section from this skill]
Extraction rules:
[Paste the Extraction Protocol section from this skill]
Apply these errata corrections: [LIST]
Save output to: [OUTPUT_PATH]
After all agents complete, verify in the main conversation and build the synthesis files.
For single-PDF books
Use the Read tool with page ranges: pages: "1-20", pages: "21-40", etc. Process sequentially or launch agents with page ranges.
Quality Standards
Mathematical Precision
- Theorem conditions: NEVER omit conditions. A theorem without its conditions is wrong. If the theorem states "for all $\gamma \in [0, 1)$", include that.
- Notation consistency: Use the book's notation throughout. Don't switch between $v(s)$ and $V(s)$ unless the book does.
- LaTeX accuracy: Verify subscripts, superscripts, summation bounds, and matrix notation. PDF extraction commonly mangles these.
Structural Fidelity
- Section mapping: Every section heading
N.Min the book gets a## N.Mheading in the markdown - No reorganization: Don't move content between sections. If the book puts Example 3.2 in Section 3.4, keep it there.
- Preserve sequence: Definitions → theorems → proofs → examples → remarks, in the order they appear
Completeness Targets
| Book Size | Expected Output | Lines per Chapter | |-----------|----------------|-------------------| | < 150 pages | 3,000-5,000 lines total | 200-400 | | 150-300 pages | 5,000-8,000 lines total | 400-700 | | 300-500 pages | 8,000-12,000 lines total | 500-800 | | 500+ pages | 12,000+ lines total | 500-800 |
Common Failure Modes
1. PDF text extraction garbles math
Symptom: Subscripts lost, summation symbols become "P", matrices flatten to single lines. Fix: Cross-check extracted text against the PDF visually. For heavily formatted math, read the PDF page as an image and transcribe manually. Consider using lecture slides as a secondary source — they often have cleaner text extraction.
2. Sub-agent can't read PDFs
Symptom: Agent reports "file not found" or empty extraction. Fix: Verify PDF reading works BEFORE launching agents (Phase 0). If the Read tool fails, pre-extract to text files:
uv run --with pymupdf python3 -c "
import fitz
doc = fitz.open('book.pdf')
for i in range(len(doc)):
with open(f'/tmp/page-{i+1:03d}.txt', 'w') as f:
f.write(doc[i].get_text())
"
Then give agents the text files instead of the PDF.
3. Chapter exceeds agent context window
Symptom: Agent output truncates mid-chapter, missing later sections. Fix: Split into 20-page chunks. Give each chunk to a separate agent, then merge results. Alternatively, process the chapter in the main conversation where context is managed automatically.
4. Theorem conditions are incomplete
Symptom: Theorem states a result but the conditions under which it holds are vague or missing. Fix: This is the most dangerous failure mode. Cross-check every theorem against the PDF. If conditions are unclear, check the errata and lecture slides. Mark uncertain conditions with [VERIFY].
5. Equation numbering drifts
Symptom: Equation (7.5) in the markdown is actually (7.6) in the book. Fix: After completing a chapter, do a sequential pass comparing equation tags against the PDF. This is fastest as a visual scan.
6. Cross-reference index is stale
Symptom: Index references sections that were renamed or reorganized during editing. Fix: Build the cross-reference index LAST, after all chapter files are finalized. Generate it by scanning the actual files, not from memory.
Worked Example: Reference Output
The knowledge base at Book-Mathematical-Foundation-of-Reinforcement-Learning/knowledge_base/ is the reference implementation of this skill. It converts a 10-chapter, 270-page RL textbook into:
- 13 files, ~7,900 lines total (~45K tokens)
- Chapter files range from 355 lines (Ch 1, foundational definitions) to 838 lines (Ch 8, most algorithms)
- Overview file: 180 lines with dependency diagram, chapter summary table, concept flow, navigation tips
- Cross-reference index: 323 lines with concept index (80+ entries), algorithm comparison (15 algorithms), theorem index (20+ theorems), equation quick-reference (~20 equations)
Use this as a template for structure, formatting, and level of detail.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: chuongdlb
- Source: chuongdlb/agent-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.