AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Mineru Extract

skill-blessonism-openclaw-skills-mineru-extract · by blessonism

Use the official MinerU (mineru.net) parsing API to convert a URL (HTML pages like WeChat articles, or direct PDF/Office/image links) into clean Markdown + structured outputs. Use when web_fetch/browser can’t access or extracts messy content, and you want higher-fidelity parsing (layout/table/formula/OCR).

No reviews yet
0 installs
27 views
0.0% view→install

Install

$ agentstack add skill-blessonism-openclaw-skills-mineru-extract

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-blessonism-openclaw-skills-mineru-extract)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
6mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Mineru Extract? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

MinerU Extract (official API)

Use MinerU as an upstream “content normalizer”: submit a URL to MinerU, poll for completion, download the result zip, and extract the main Markdown.

Quick start (MCP-aligned)

We align to the MinerU MCP mental model, but we do not run an MCP server.

  • Primary script (MCP-style): scripts/mineru_parse_documents.py
  • Input: --file-sources (comma/newline-separated)
  • Output: JSON contract on stdout: { ok, items, errors }
  • Low-level script (single URL): scripts/mineru_extract.py

Auth:

  • Set MINERU_TOKEN (Bearer token from mineru.net)

Default model heuristic:

  • URLs ending with .pdf/.doc/.ppt/.png/.jpgpipeline
  • Otherwise → MinerU-HTML (best for HTML pages like WeChat articles)

1) Configure token (skill-local)

Put secrets in skill root .env (do not paste into chat outputs):

# /home/node/.openclaw/workspace/skills/mineru-extract/.env
MINERU_TOKEN=***
MINERU_API_BASE=https://mineru.net

2) Parse URL(s) → Markdown (recommended)

MCP-style wrapper (returns JSON, optionally includes markdown text):

python3 skills/mineru-extract/scripts/mineru_parse_documents.py \
  --file-sources "\n" \
  --language ch \
  --enable-ocr \
  --model-version MinerU-HTML

If you want the markdown content inline in the JSON (can be large):

python3 skills/mineru-extract/scripts/mineru_parse_documents.py \
  --file-sources "" \
  --model-version MinerU-HTML \
  --emit-markdown --max-chars 20000

Low-level (single URL, print markdown to stdout):

python3 skills/mineru-extract/scripts/mineru_extract.py "" --model MinerU-HTML --print > /tmp/out.md

Output

The script always downloads + extracts the MinerU result zip to:

/home/node/.openclaw/workspace/mineru//

It writes:

  • result.zip
  • extracted files (Markdown + JSON + assets)

It prints a JSON summary to stderr with paths:

  • task_id, full_zip_url, out_dir, markdown_path

Parameters (common)

  • --model: pipeline | vlm | MinerU-HTML (HTML requires MinerU-HTML)
  • --ocr/--no-ocr: enable OCR (effective for pipeline/vlm)
  • --table/--no-table: table recognition
  • --formula/--no-formula: formula recognition
  • --language ch|en|...
  • --page-ranges "2,4-6" (non-HTML)
  • --timeout 600 / --poll-interval 2

Failure modes & fallbacks

  • MinerU may fail to fetch some URLs (anti-bot / geo / login).
  • Fallback: provide an HTML file or a PDF/long screenshot; then implement “upload + parse” flow with MinerU batch upload endpoints.
  • Always report the failing URL + MinerU err_msg and keep an original-source link in outputs.

References

  • MinerU API docs: https://mineru.net/apiManage/docs
  • MinerU output files: https://opendatalab.github.io/MinerU/reference/output_files/

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.