# Docx Shell Extract

> Extract text from DOCX files using shell commands when python-docx is unavailable

- **Type:** Skill
- **Install:** `agentstack add skill-hkuds-openspace-docx-shell-extract`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [HKUDS](https://agentstack.voostack.com/s/hkuds)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [HKUDS](https://github.com/HKUDS)
- **Source:** https://github.com/HKUDS/OpenSpace/tree/main/gdpval_bench/skills/docx-shell-extract

## Install

```sh
agentstack add skill-hkuds-openspace-docx-shell-extract
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# DOCX Shell Extraction

## When to Use This Skill

Use this pattern when you need to read or extract text from Microsoft Word (.docx) files in constrained environments where:
- The `python-docx` library is not available
- You cannot install additional Python packages
- You need a quick, reliable shell-based solution

## Core Technique

DOCX files are ZIP archives containing XML files. The main document content is stored in `word/document.xml`. You can extract and parse this using standard shell tools.

## Step-by-Step Instructions

### Step 1: Extract the document.xml content

```bash
unzip -p filename.docx word/document.xml
```

The `-p` flag pipes the content to stdout without extracting to disk.

### Step 2: Strip XML tags to get plain text

```bash
unzip -p filename.docx word/document.xml | sed 's/]*>//g'
```

This removes all XML tags, leaving the text content.

### Step 3: Clean up whitespace (optional)

For cleaner output, add additional sed processing:

```bash
unzip -p filename.docx word/document.xml | \
  sed 's/]*>//g' | \
  sed 's/&[^;]*;//g' | \
  sed 's/^[[:space:]]*//' | \
  sed 's/[[:space:]]*$//' | \
  sed '/^$/d'
```

This removes:
- XML tags
- XML entities (like `&amp;`, `&lt;`)
- Leading/trailing whitespace
- Empty lines

### Step 4: Save to a text file (optional)

```bash
unzip -p filename.docx word/document.xml | \
  sed 's/]*>//g' > output.txt
```

## Complete Example

```bash
# Extract text from a Word document
DOCX_FILE="report.docx"
OUTPUT_FILE="report_text.txt"

unzip -p "$DOCX_FILE" word/document.xml | \
  sed 's/]*>//g' | \
  sed 's/&[^;]*;//g' | \
  sed '/^$/d' > "$OUTPUT_FILE"

echo "Extracted text saved to $OUTPUT_FILE"
```

## Verification

After extraction, verify the content was captured:

```bash
# Check if output file has content
if [ -s "$OUTPUT_FILE" ]; then
    echo "Successfully extracted $(wc -l < "$OUTPUT_FILE") lines"
    head -5 "$OUTPUT_FILE"
else
    echo "Warning: Output file is empty"
fi
```

## Limitations

- This method extracts raw text without formatting
- Complex layouts, tables, and images are not preserved
- Some special characters may need additional handling
- Works best for text-heavy documents

## Alternatives to Explore

If this approach fails or the DOCX structure differs:
- Check for `word/document.xml` existence: `unzip -l filename.docx | grep document.xml`
- Some documents may use `word/*.xml` with different naming
- Consider `pandoc` if available: `pandoc filename.docx -t plain`

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [HKUDS](https://github.com/HKUDS)
- **Source:** [HKUDS/OpenSpace](https://github.com/HKUDS/OpenSpace)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-hkuds-openspace-docx-shell-extract
- Seller: https://agentstack.voostack.com/s/hkuds
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
