AgentStack
MCP verified MIT Self-run

Datoon

mcp-andrii-su-datoon · by andrii-su

Smart JSON-to-TOON conversion with pragmatic auto-gating for LLM prompts

No reviews yet
0 installs
6 views
0.0% view→install

Install

$ agentstack add mcp-andrii-su-datoon

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access Used
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Datoon? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

datoon

smart structured-data→TOON gateway — converts only when it actually saves tokens

Before/After • Install • What You Get • How It Works • Benchmarks • Full install guide

______________________________________________________________________

Raw structured data is often verbose in LLM prompts. TOON can save tokens — but blind conversion can also make payloads worse. datoon adds a decision layer: convert when structure and savings justify it, skip when they don't, and always explain why.

Supports JSON, CSV, JSONL, YAML, XML, Parquet, Avro, ORC, Excel, and Apple Numbers — auto-detected from file extension.

Before / After

JSON in the prompt (43 tokens)

{"users":[
  {"id":1,"name":"Ada","role":"admin"},
  {"id":2,"name":"Lin","role":"analyst"},
  {"id":3,"name":"Grace","role":"viewer"}
]}

datoon converts → TOON (24 tokens)

users[3]{id,name,role}:
  1,Ada,admin
  2,Lin,analyst
  3,Grace,viewer
{"decision":"convert","reason":"Estimated savings 44.19% (threshold 15.00%)."}

CSV from a data pipeline (111 tokens as JSON)

id,name,role
1,Ada,admin
2,Lin,analyst
3,Grace,viewer

datoon auto-converts → TOON (24 tokens)

datoon data.csv --report-stdout

Same result. Zero JSON serialization in your code.

Non-uniform payload (26 tokens)

{"config":{"debug":true},"tags":["a","b"]}

datoon skips → keeps JSON

{"decision":"skip","reason":"No uniform object arrays found with at least 3 rows."}

No Node.js call. No silent corruption.

Same data. Right format. Always explained.

┌──────────────────────────────────────────────────┐
│  PAYLOAD SAVINGS (auto avg)    ████░░░░░░   28%  │
│  PAYLOAD SAVINGS (agent skill) ████████░░   62%  │
│  DECISION ACCURACY             ██████████  100%  │
│  HARMFUL CONVERSIONS BLOCKED   ██████████  100%  │
└──────────────────────────────────────────────────┘

> [!IMPORTANT] > datoon saves payload tokens — the structured data portion of your prompt. Token savings depend on payload shape: uniform tabular data converts well; deeply nested or non-uniform structures are skipped. Every decision includes a reason so pipelines can log, debug, and trust the outcome.

Install

# core (JSON, CSV, JSONL, XML — no extra deps)
uv add datoon
pip install datoon

# with YAML support
pip install "datoon[yaml]"

# with Excel support
pip install "datoon[excel]"

# with Parquet / ORC / Avro support
pip install "datoon[columnar]"

# with Apple Numbers support
pip install "datoon[numbers]"

# with tiktoken-based token counting
pip install "datoon[tokens]"

# with MCP server
pip install "datoon[mcp]"

# everything
pip install "datoon[all]"

Requires Python 3.12+. TOON conversion requires Node.js with npx in PATH — analysis and format reading work without it.

For Claude Code plugin, Codex, and MCP config → [INSTALL.md](./INSTALL.md).

What You Get

| | What | |---|---| | datoon CLI | Auto-gate any supported format → TOON from terminal or scripts | | Python API | convert_json_for_llm() + read_tabular() for any LLM pipeline | | MCP Server | convert_json, convert_text, analyze_json tools for Claude Desktop, Cursor, Windsurf | | Claude Code Plugin | /datoon in-session trigger, installs from GitHub in one command | | Codex Plugin | Marketplace plugin — structured-data mode for Codex |

Supported input formats

| Format | Extension | Extra needed | |---|---|---| | JSON | .json | — | | JSONL | .jsonl, .ndjson | — | | CSV | .csv | — | | XML | .xml | — | | YAML | .yaml, .yml | datoon[yaml] | | Excel | .xlsx, .xls | datoon[excel] | | Parquet | .parquet | datoon[columnar] | | Avro | .avro | datoon[columnar] | | ORC | .orc | datoon[columnar] | | Apple Numbers | .numbers | datoon[numbers] |

How It Works

  1. Detect format — from --format flag, file extension, or default to JSON for stdin
  2. Read + normalize — parse source into list of row dicts; serialize to compact JSON
  3. Analyze structure — uniform object arrays? acceptable depth? minimum rows?
  4. Gate early — non-candidates skip before any CLI call; no Node.js overhead
  5. Convert + estimate — TOON CLI runs, token savings calculated
  6. Gate savings — below threshold → return JSON; above → return TOON with report

Every path returns a ConversionReport with decision, reason, and token estimates. Pipelines never get silent surprises.

______________________________________________________________________

Quick Start

JSON (stdin):

echo '{"users":[{"id":1,"name":"Ada"},{"id":2,"name":"Lin"},{"id":3,"name":"Grace"}]}' | datoon --report-stdout

CSV (auto-detected from extension):

datoon data.csv --report-stdout

JSONL:

datoon data.jsonl -o output.toon

YAML (requires datoon[yaml]):

datoon data.yaml --report-stdout

Parquet (requires datoon[columnar]):

datoon data.parquet --report ./report.json

Explicit format override:

datoon --format csv ` | — | Write JSON conversion report to file |
| `--report-stdout` | — | Print JSON conversion report to stderr |
| `-o ` | stdout | Output file path |
| `--version` | — | Print version and exit |

Format is auto-detected from file extension. Use `--format` to override or when reading from stdin.

______________________________________________________________________

## Benchmarks

```bash
PYTHONPATH=src python benchmarks/run.py --dry-run
PYTHONPATH=src python benchmarks/run.py
PYTHONPATH=src python benchmarks/run.py --update-readme

Why auto mode outperforms forced conversion

Auto mode avoids low-benefit and high-risk payloads (orders-nested, mixed-non-uniform) while matching forced TOON's average token count on suitable ones. Every decision comes with a reasoned report.

| Scenario | JSON Baseline | Forced TOON | datoon Auto | |---|---:|---:|---:| | Average tokens | 77 | 50 | 50 | | Avg token saved | 0.0% | 26.8% | 28.1% | | Decision quality | n/a | Converts all | Converts 3/5, skips harmful cases |

| Dataset | JSON | TOON (forced) | Raw Saved | Auto | Auto Tokens | Auto Saved | |---|---:|---:|---:|---|---:|---:| | users-small | 54 | 40 | 25.9% | convert | 40 | 25.9% | | events-medium | 219 | 162 | 26.0% | convert | 162 | 26.0% | | orders-nested | 106 | 116 | -9.4% | skip | 106 | 0.0% | | mixed-non-uniform | 35 | 47 | -34.3% | skip | 35 | 0.0% | | metrics-wide | 142 | 103 | 27.5% | convert | 103 | 27.5% | | Average | 111 | 94 | 7.1% | 3/5 convert | 89 | 15.9% |

Forced conversion succeeded for 5/5 payloads.

Format conversion benchmark

Token savings when converting from common structured formats (CSV, JSONL, XML, YAML). Baseline is the JSON representation of the same data — what an LLM would receive without datoon.

| Dataset | Format | JSON Tokens | TOON (forced) | Auto | Auto Tokens | Auto Saved | |---|---|---:|---:|---|---:|---:| | users-csv | csv | 53 | 29 | convert | 29 | 45.3% | | events-jsonl | jsonl | 194 | 109 | convert | 109 | 43.8% | | catalog-xml | xml | 96 | 50 | convert | 50 | 47.9% | | metrics-yaml | yaml | 129 | 61 | convert | 61 | 52.7% | | Average | — | 118 | 62 | 4/4 convert | 62 | 47.4% |

Forced conversion succeeded for 4/4 payloads.

Agent skill evaluation

Artifact-based subagent comparison — identical analysis tasks, two modes:

  • with_skill: agent received the datoon skill and followed the conversion workflow.
  • without_skill: agent used JSON directly, no TOON or datoon.

3 payload sizes × 3 iterations = 18 total agent runs. Both modes: 100% correct answers.

| Scenario | Avg JSON Tokens | Avg TOON Tokens | Avg Payload Saved | |---|---:|---:|---:| | small | 225 | 118 | 47.6% | | medium | 2,972 | 1,138 | 61.7% | | large | 17,757 | 6,673 | 62.4% |

Full report and raw outputs: [benchmarks/agent_skill_eval/](benchmarks/agentskilleval/). Savings are payload-token estimates, not full end-to-end model-token usage.

______________________________________________________________________

Development

Contributor workflow: [CONTRIBUTING.md](./CONTRIBUTING.md). Maintainer/agent notes: [CLAUDE.md](./CLAUDE.md).

Setup:

uv sync --extra dev
uvx pre-commit install

Tests:

pytest -m "not integration"   # unit only (102 tests)
pytest                        # with integration (requires Node.js + npx)

Skill sync + plugin metadata:

python scripts/validate_skill_sync.py
python scripts/validate_plugin_metadata.py

______________________________________________________________________

Links

  • [INSTALL.md](./INSTALL.md) — full install matrix, all targets, per-agent detail
  • [CONTRIBUTING.md](./CONTRIBUTING.md) — contributor workflow
  • [CLAUDE.md](./CLAUDE.md) — maintainer guide for agents
  • [CHANGELOG.md](./CHANGELOG.md) — release history
  • [SECURITY.md](./SECURITY.md) — vulnerability reporting
  • Live docsdocs/
  • Issues — bugs, features, questions

______________________________________________________________________

License

MIT

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.