Install
$ agentstack add skill-chuongdlb-agent-skills-tex-source-paper-extractor ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
TeX-Source Paper Extractor
Purpose
Produce or enrich a structured "paper card" by reading a paper's arXiv LaTeX e-print source rather than its compiled PDF. The card format and evaluative-extraction philosophy are identical to paper-extractor — this skill only changes the input pathway. TeX source is preferred for arXiv papers because:
- Formulas are already in LaTeX (no lossy PDF math OCR) — satisfies the KB rule that all formulas stay as
$inline$/$$display$$. - Tables, algorithms, and appendices (hyperparameters, ablations, proofs) are cleanly accessible — these are exactly the parts PDF/metadata extraction tends to drop.
- The e-print is often a newer/extended version (e.g. journal/IJRR) than the originally-cited venue.
When to Use
- Creating a new card for a paper that is on arXiv.
- Enriching an existing thin or PDF-built card for an arXiv paper (the common case — see the
tex-enrichment-batches.mdinitiative in the papers KB). - Any time formula fidelity matters.
Not for: non-arXiv papers (fall back to paper-extractor on the PDF), lossless book extraction (book-to-knowledge-base), or topic synthesis (kb-integrator).
Inputs
- arXiv ID (e.g.
2303.04137) and the target card id/path. - KB taxonomy at
kb/config/taxonomy.mdfor domain tags.
Procedure
1. Download the e-print source
curl -sL -A "Mozilla/5.0 (research KB)" "https://arxiv.org/e-print/" -o src.tar.gz
tar xzf src.tar.gz 2>/dev/null || gunzip -c src.tar.gz > main.tex # some are a single gzipped .tex
The endpoint returns a gzipped tarball (or, rarely, a single gzipped .tex).
2. Find the main file and follow includes
grep -rl "\\begin{document}" --include="*.tex" .
Multi-file papers use \input{...} / \include{...} — read the included section files (often under sections/, text/). Appendices and supplementary (supp.tex, appendix.tex) carry the highest-value extra content.
3. Extract following the paper-extractor card format
Use the same card template (frontmatter + One-Line Summary, Problem, Contributions, Method Summary, Key Results table, Baselines, Limitations, Novelty Claims, Relevance). Then specifically harvest from TeX what PDFs lose:
- Exact equations (copy LaTeX verbatim).
- Hyperparameter tables and training details (often appendix).
- Ablation results and secondary findings.
- Theoretical derivations / connections (e.g. control-theory limits, EBM/score-function arguments).
4. Resolve custom LaTeX macros — CRITICAL
Papers define private macros (\newcommand{\obs}{...}, \vox, \shortname, \ours, \qattn, …). These will not render in MathJax/markdown. Before writing the card:
- Check
\newcommand/\defdefinitions (often inmacros.tex,notation.tex, or the preamble). - Replace each macro with standard renderable notation (e.g.
\obs → o,\vox → V). - Verify no leftover macros:
grep -E '\\(obs|vox|shortname|ours|qattn)\b'.
5. Verify metadata against the source
Stub/old cards frequently have wrong metadata. Fix from the TeX \author{}, title, and any venue macro:
authors: ["Various"]→ real author list.- Wrong/placeholder venue (e.g. arXiv-format default) → correct venue.
- Version drift: if the e-print is an extended/journal version (look for journal class files like
iclr2025_conference,IEEEtran,ijrrfigure prefixes,\iclrfinalcopy), do not silently attribute later-version content to the originally-cited venue. Note the mismatch.
6. Save and quality-check
Write to kb/papers/.md. Run the paper-extractor quality checklist plus: all formulas LaTeX, no unresolved macros, metadata matches source.
Batch / token-efficiency mode (gap scan)
When enriching many already-rich cards, don't read every full TeX into context. Use a mechanical gap-scan first to rank which cards actually have missing content (high word-ratio is expected and not itself a defect — look for appendices, many equations/tables, and substantive uncovered section titles). The reusable script and the 13-batch plan live in the papers KB at kb/reports/tex-enrichment-batches.md.
Common Pitfalls
- Leftover macros rendering as raw
\obsin the card — always resolve (step 4). - Missing
\inputfiles — the main.texmay be near-empty; the content is in included section files. - Treating the e-print as the cited version — it may be newer; check class files (step 5).
- Over-enriching — cards are evaluative summaries, not transcriptions; pull what's KB-relevant (method/results/theory), skip acknowledgements, author lists, reproducibility boilerplate.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: chuongdlb
- Source: chuongdlb/agent-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.