# Meta Data Mcp

> A meta-mcp (mmcp) server offering a window into the exciting world of open data sources by intelligently routing model tool calls through an ever-growing index of open-data sources.

- **Type:** MCP server
- **Install:** `agentstack add mcp-derekslinz-meta-data-mcp`
- **Verified:** Pending review
- **Seller:** [derekslinz](https://agentstack.voostack.com/s/derekslinz)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [derekslinz](https://github.com/derekslinz)
- **Source:** https://github.com/derekslinz/meta-data-mcp
- **Website:** https://www.linzalytics.com

## Install

```sh
agentstack add mcp-derekslinz-meta-data-mcp
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# meta-data-mcp

> A single MCP server that transparently routes user requests to 88 open-data sources.

`meta-data-mcp` is one MCP server — not many. Under the hood it bundles 88 *plugins*, each wrapping a different open-data API. The plugins are an implementation detail; from your LLM's perspective there is one server and one place to ask "where can I find data about X?"

You install one server. You get all the data, discoverable through built-in routing tools.

## Why "meta"?

Finding open data isn't the hard part — there's an absurd amount of it available. The hard part is finding the right dataset *when you need it*. `meta-data-mcp` makes that automatic:

- The LLM calls `opendata_providers_find` ("FX rates", "court rulings", "earthquakes near Lisbon") and the server routes the query against an internal registry of every bundled plugin.
- The LLM then calls the matching tool directly. No setup step in between, no separate servers, no per-provider install rituals.

This project was forked from [opendata-mcp](https://github.com/OpenDataMCP/OpenDataMCP) and reshaped around the single-server idea once the catalogue passed a few dozen plugins.

## Installation

You'll need `uv` (a Python package manager).

```bash
# macOS — install uv via Homebrew so MCP clients can find it
brew install uv

# Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
```

Then register the server with every MCP client installed on your machine:

```bash
uv run meta-data-mcp setup
```

The command auto-detects which MCP clients you have installed and adds **one** `meta-data-mcp` entry under `mcpServers` in each. Supported clients:

| Client | Config file |
|---|---|
| Claude Desktop | `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS) / `%APPDATA%/Claude/claude_desktop_config.json` (Windows) |
| Claude Code | `~/.claude.json` |
| Cursor | `~/.cursor/mcp.json` |
| Windsurf | `~/.codeium/windsurf/mcp_config.json` |
| Gemini CLI | `~/.gemini/settings.json` |
| LM Studio | `~/.cache/lm-studio/mcp.json` |

Each existing config is backed up to `.bak` before writing. Restart the affected client(s) and you'll see one new server with discovery tools available immediately; plugin tools can then be activated on demand.

Inspect what's detected / configured on your machine:

```bash
uv run meta-data-mcp clients
```

Target a single client (or write to every supported client regardless of detection):

```bash
uv run meta-data-mcp setup --client claude-code
uv run meta-data-mcp setup --client all
```

If you want to see the JSON snippet without touching any config file (e.g. to paste into a client we don't support yet):

```bash
uv run meta-data-mcp setup --print-json
```

When `META_DATA_MCP_AUTH_TOKEN` is set, `--print-json` also surfaces the SSE-client snippet (with the real token) to stderr so you can wire a remote client.

### Hosting `meta-data-mcp` as a remote SSE server

For deploying behind your own domain with bearer-token authentication, see [`docs/hosting.md`](docs/hosting.md). It covers `systemd`, Caddy/nginx TLS termination, token rotation, and the threat model.

## CLI

There is one server, so the CLI takes no "provider" argument. Every command operates on the one `meta-data-mcp` server.

| Command | What it does |
|---|---|
| `uv run meta-data-mcp run` | Run the server (default SSE; pass `--transport stdio` for Claude Desktop). |
| `uv run meta-data-mcp setup` | Register the server in detected MCP client configs (or one target via `--client`). |
| `uv run meta-data-mcp remove` | Unregister the server from detected MCP client configs (or one target via `--client`). |
| `uv run meta-data-mcp cleanup` | Detect and remove legacy multi-server entries (`--apply` to commit). |
| `uv run meta-data-mcp inspect` | Launch [mcp-inspector](https://modelcontextprotocol.io/docs/tools/inspector) against the server. |
| `uv run meta-data-mcp list` | Informational: list the internal plugins bundled in this server. |
| `uv run meta-data-mcp info` | Informational: show server overview. Pass `--plugin ` for plugin-level details. |
| `uv run meta-data-mcp version` | Print the package version. |

The `list` command exists for transparency about what's bundled — **plugins are not separately installable, runnable, or addressable**. They are loaded automatically when the server starts.

## Server tools (what the LLM calls)

Once `meta-data-mcp` is running, the LLM has access to two layers of tools — and you don't need to mention either to the user:

1. **Meta tools** — the 13 server-level tools below. They make routing transparent: the LLM uses them to find, activate, and (if needed) create the right plugin without you telling it which tool to call.
2. **Plugin tools** — ~330 tools coming from the 88 bundled plugins. In the default discovery-only mode they are activated per provider at runtime (or preloaded via `META_DATA_MCP_PRELOAD`). The LLM picks one after consulting the meta tools.

### Meta tools

| Tool | Purpose |
|---|---|
| `opendata_providers_find` | Free-text search over the plugin registry. Returns ranked matches. When nothing matches the response carries a `no_match: true` flag and a `next_step` hint pointing at `opendata_plugins_draft` + `opendata_plugins_create`. |
| `opendata_explain_choice` | Show the scoring breakdown for a search (useful for debugging routing decisions). |
| `opendata_domains_list` | Enumerate the controlled domain vocabulary (`health`, `legal`, `finance`, `earth-science`, …). |
| `opendata_regions_list` | Enumerate the controlled region vocabulary (`us`, `eu`, `uk`, `global`, …). |
| `opendata_providers_describe` | Full metadata for one plugin by id — title, description, domains, regions, keywords, homepage, required env vars. |
| `opendata_providers_list` | Paginated dump of the whole registry. |
| `opendata_providers_activate` | Activate one provider so its tools become callable in this session. |
| `opendata_providers_deactivate` | Remove an activated provider's tools from the current session catalog. |
| `opendata_providers_list_active` | List currently active providers and the tool names each contributes. |
| `opendata_health_snapshot` | Return per-provider health scores used by discovery health badges and routing context. |
| `opendata_plugins_draft` | **Build a validated plugin YAML spec from structured inputs.** Takes id, base_url, tool definitions (name, endpoint, params), and registry metadata. Validates id/tool-name casing, path-placeholder/param consistency, and parameter types, then emits a YAML string ready to feed into `opendata_plugins_create`. Use this so the LLM never has to hand-author YAML. |
| `opendata_plugins_create` | **Autonomously create a new plugin.** Takes a YAML spec (typically produced by `opendata_plugins_draft`), runs the generator, imports the new module, registers it in the live registry, and hot-loads its tools onto the running server. Use this when `opendata_providers_find` returns no match. |
| `opendata_tool_call` | Proxy-call an activated plugin tool by name for environments that cannot directly invoke dynamically added tools. |

### The autonomous discovery flow

The reason this server is called "meta" is that it routes data requests on the user's behalf — including by *creating* the route when one doesn't exist yet. The full flow:

1. **User asks for data**, e.g. "show me the most recent published CVEs."
2. **LLM calls `opendata_providers_find`** with the query (`cve`, `vulnerability`, …).
3. **If the registry has a match**: the LLM activates the matching provider (`opendata_providers_activate`, or `activate_top` in find) and then calls the plugin tool.
4. **If the registry has no match**: the response includes `no_match: true` and a `next_step` field that explains the autonomous creation path. The LLM:
   1. Tells the user it's about to add coverage for this data source.
   2. Web-searches for an open API that exposes the requested data (e.g. the NVD or CIRCL CVE API).
   3. Calls `opendata_plugins_draft` with the API's id, base URL, and structured tool definitions. The server validates the inputs (id casing, path-placeholder consistency, parameter types) and returns a YAML string.
   4. Passes that YAML to `opendata_plugins_create`. The server materializes the plugin module + tests, imports the module, registers a `ProviderEntry` in the in-memory dynamic registry, and merges the new tools into the running server's tool list.
   5. Calls the newly-available tool to answer the user's original question.
5. **User gets their answer** — and the plugin remains available for the rest of the session.

The materialized plugin lives on disk (`meta_data_mcp/providers/{id}.py` + `tests/providers/test_{id}.py`); contributors can clean it up, add it to `meta_data_mcp/registry.py` as a static entry, and open a PR so it becomes part of every shipped install.

### Plugin tools

Every bundled plugin contributes its own tools under the one server. Their names are unique kebab-case identifiers, often using a provider-specific prefix (e.g. `usgs-eq-feed-significant-week`, `frankfurter-latest`, `wikipedia-fetch-summary`). The LLM discovers them through `opendata_providers_find`/`opendata_providers_describe`, activates the provider when needed, and can inspect session state with `opendata_providers_list_active`.

## Presentation layer (MCP Apps)

v2.0 adds a visual layer on top of every tool result. Hosts that support the [MCP Apps extension](https://modelcontextprotocol.io/docs/extensions/apps) (Claude Desktop, MCP Inspector, others) render bound tool results inline as interactive panels in a sandboxed iframe instead of as JSON text. Hosts that don't speak MCP Apps fall back to the same JSON they always got — the binding is purely additive.

Each MCP-Apps-aware tool declares its panel via `_meta.ui.resourceUri` on the tool description. The host fetches the `ui://` resource (HTML + bundled JS, single payload, no external requests besides explicitly-whitelisted CDNs) and dispatches bidirectional `postMessage` events between the iframe and itself.

### Shape primitives — `ui://meta-data-mcp/shape//v1`

Three reusable bundles cover the common payload contracts. Any tool whose response matches one of these shapes binds to the corresponding primitive automatically and gets a rich renderer for free.

| Shape | Renders | Payload contract |
|---|---|---|
| `timeseries/v1` | Line chart + auto-computed profile (min/max/mean/stddev/gap-count) via Plotly. | `{points: [{date, value, series?}], axes: {x, y}, annotations?}` |
| `geofeatures/v1` | Leaflet map + marker cluster (with density layer for high-cardinality outputs). | `{features: GeoJSON | [{lat, lon, attrs}], layers?, facets?}` |
| `records/v1` | Faceted, sortable, paginated HTML table + per-column auto-profile (type inference, top-k, null rate, range). | `{rows: [...], schema?, default_facets?}` |

### Custom apps — `ui://meta-data-mcp/app//v1`

Some data shapes don't fit a generic primitive. v2.0 ships dedicated apps for them:

| App | Drives | Visualization |
|---|---|---|
| `discovery/v1` | `opendata_providers_find`, `opendata_domains_list`, `opendata_regions_list`, `opendata_providers_activate`, etc. | Faceted plugin browser with live health badges. |
| `vulnerability/v1` | `nvd-*`, `osv-*`, `epss-*`, `cisa-kev`. | CVSS radar + severity heatmap + exploitation-probability gauge. |
| `entity-graph/v1` | `crossref-works-by-author`, `openalex-search-works`, `wikidata-search-entities`, `opensanctions-search`. | Force-directed graph (D3) with co-authorship overlay. |
| `trade-flows/v1` | `comtrade-trade-data`. | Reporter → commodity → partner Sankey + commodity treemap. |
| `news-tone/v1` | `gdelt-article-search`, `gdelt-volume-timeline`. | Volume + tone timeline with country-pair chord diagram. |
| `network-topology/v1` | `ripestat-asn-neighbours` and friends. | Force-directed ASN peering/upstream/downstream graph. |
| `molecular/v1` | `pubchem-compound`, `pdb-entry`. | WebGL 3D structure viewer (3Dmol.js, cartoon for proteins, stick+sphere for ligands). |
| `museum/v1` | `met-search`, `met-search-by-artist`, `met-get-object`. | Lazy-loaded CSS-grid image gallery + provenance detail panel. |

### Building new apps

Adding a UI binding to a generated provider is now a one-line spec change:

```yaml
tools:
  - name: my-tool
    description: ...
    endpoint: /foo
    response_shape: records   # ← binds to the shape primitive
```

See [`tools/specs/README.md`](tools/specs/README.md) for the full reference. Bundle-size budgets are enforced in CI (warn ≥ 100 KB, error ≥ 1 MB); the v2.0 bundles range from 14 KB (timeseries primitive) to 34 KB (vulnerability app), all comfortably inside the budget.

## Citable answers

Every tool result carries a machine-readable **citation manifest**: exactly which upstream requests produced it. The transport kernel records each HTTP exchange during a tool call, and the result's first content block gains a `_meta["meta-data-mcp/citations"]` entry:

```jsonc
{
  "sources": [
    {
      "provider": "eu-eurostat",
      "title": "Eurostat",
      "homepage": "https://ec.europa.eu/eurostat",
      "license": "Eurostat data is reusable under CC BY 4.0; cite '© European Union, Eurostat'.",
      "url": "https://ec.europa.eu/eurostat/api/dissemination/statistics/1.0/data/nama_10_gdp?format=JSON&lang=en",
      "method": "GET",
      "status": 200,
      "fetched_at": "2026-07-09T14:02:11.482Z",
      "cache_hit": false
    }
  ]
}
```

This is what makes an LLM data answer auditable: the exact URL(s) — query parameters included — when they were fetched, whether they came from the transport cache, and the provider's license/attribution terms. Anyone can re-issue the URL and check the claim.

- **Secrets never leak.** Values of sensitive query parameters are replaced with `REDACTED` — an exact denylist (`api_key`, `token`, `appid`, …) plus conservative heuristics (`*key`, `*token`, `*secret*`, `*signature*`, …) that also cover presigned cloud-storage URLs and plugin-specific key params. Userinfo credentials in the URL itself (`https://user:pass@host`) are redacted too; parameter names are preserved so the URL stays reproducible with your own credentials. Headers never enter the manifest.
- **Failed exchanges are cited too** — a 4xx/5xx a handler recovered from, and the intermediate 429/5xx attempts the kernel's retry loop absorbed, are part of how the answer was produced; filter on `status`. (A tool call that *errors out* returns the SDK's `isError` result, which carries no manifest.)
- **Honest timestamps.** `fetched_at` is when the bytes were actually fetched: cache-served exchanges report the original fetch time with `cache_hit: true`, not the cache-read time.
- **On by default.** Set `META_DATA_MCP_CITATIONS=0` to disable. Complements the opt-in tamper-evidence digest (`META_DATA_MCP_PROVENANCE`); both can coexist on the same result.

## Bundled plugins (88)

This is what's inside the one server. You don't install these individually — they all come along.

### Government / Civic

| Plugin | Source | Description |
|---|---|---|
| `au_data_gov` | Australian Government Open Data | CKAN catalog at data.gov.au |
| `ca_open_gov` | Canada Open Data | CKAN catalog at open.canada.ca |
| `ch_opendata_swiss` | opendata.swiss | Swiss federal open-data catalog (CKAN) |
| `de_govdata` | GovData Germany | Germany's federal open-data catalog (CKAN) |
| `fr_data_gouv` | data.gouv.fr | French government open data platform |
| `nl_tweedekamer` | Tweede Kamer | Dutch Parliament open data |
| `sg_data_gov` | Singapore Open Data | data.gov.sg datasets and collections |
| `uk_gov` | data.gov.uk | UK government CKAN catalog |
| `us_cary` | Town of Cary Open Data | Town of Cary, NC open data via Socrata — public safety, transportation, utilities, parks |
| `us_data_gov` | Data.gov | US federal government ope

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [derekslinz](https://github.com/derekslinz)
- **Source:** [derekslinz/meta-data-mcp](https://github.com/derekslinz/meta-data-mcp)
- **License:** MIT
- **Homepage:** https://www.linzalytics.com

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: flagged — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-derekslinz-meta-data-mcp
- Seller: https://agentstack.voostack.com/s/derekslinz
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
