AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Audio Transcriber

mcp-knuckles-team-audio-transcriber · by Knuckles-Team

Transcribe audio into text. Agentic AI supported through MCP Server.

No reviews yet
0 installs
28 views
0.0% view→install

Install

$ agentstack add mcp-knuckles-team-audio-transcriber

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-knuckles-team-audio-transcriber)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Audio Transcriber? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Audio Transcriber

CLI or API | MCP | Agent

Version: 1.0.1

> Documentation — Installation, deployment, and usage across the CLI, Python API, > MCP server, and A2A agent are maintained in the > official documentation.


Overview

Audio Transcriber is a production-grade Agent and Model Context Protocol (MCP) server designed to interface directly with Transcribe your .wav .mp4 .mp3 .flac files to text or record your own audio!.


Key Features

  • Consolidated Action-Routed MCP Tools: Minimizes token overhead and eliminates tool bloat in LLM contexts by grouping methods into optimized, togglable tool modules.
  • Enterprise-Grade Security: Comprehensive support for Eunomia policies, OIDC token delegation, and granular execution context tracking.
  • Integrated Graph Agent: Built-in Pydantic AI agent supporting the Agent Control Protocol (ACP) and standard Web interfaces (AG-UI).
  • Native Telemetry & Tracing: Out-of-the-box OpenTelemetry exports and native Langfuse tracing.

CLI or API

This agent wraps the Transcribe your .wav .mp4 .mp3 .flac files to text or record your own audio! API. You can interact with it programmatically or via its integrated execution entrypoints.

Detailed instructions on how to use the underlying API wrappers, extended schema bindings, and developer SDK references are maintained in [docs/index.md](docs/index.md).


MCP

This server utilizes dynamic Action-Routed tools to optimize token overhead and maximize IDE compatibility.

Available MCP Tools

The table below is auto-generated from the live server — do not edit by hand.

Condensed action-routed tools (default — MCP_TOOL_MODE=condensed)

| MCP Tool | Toggle Env Var | Description | |----------|----------------|-------------| | health_check | MISCTOOL | | | transcribe_audio | AUDIO_PROCESSINGTOOL | Transcribes audio from a provided file or by recording from the microphone. |

Verbose 1:1 API-mapped tools (MCP_TOOL_MODE=verbose or both)

7 per-operation tools — one per public API method (click to expand)

| MCP Tool | Toggle Env Var | Description | |----------|----------------|-------------| | audio_transcriber_export | AUDIO_TRANSCRIBERTOOL | Export transcription to specified formats. | | audio_transcriber_initiate_stream | AUDIO_TRANSCRIBERTOOL | Initiate the audio input stream. | | audio_transcriber_interact | AUDIO_TRANSCRIBERTOOL | Interact with PersonaPlex server via WebSocket. | | audio_transcriber_record | AUDIO_TRANSCRIBERTOOL | Record audio for a specified duration or until stopped. | | audio_transcriber_save_stream | AUDIO_TRANSCRIBERTOOL | Save the recorded frames to a WAV file. | | audio_transcriber_stop_stream | AUDIO_TRANSCRIBERTOOL | Stop and close the audio stream. | | audio_transcriber_transcribe | AUDIO_TRANSCRIBERTOOL | Transcribe the audio file using the initialized backend. |

2 action-routed tool(s) (default) · 7 verbose 1:1 tool(s). Each is enabled unless its TOOL toggle is set false; MCP_TOOL_MODE selects the surface (condensed default · verbose 1:1 · both). Auto-generated — do not edit.

Detailed tool schemas, parameter shapes, and validation constraints are preserved in [docs/mcp.md](docs/mcp.md).

Dynamic Tool Selection & Visibility

This MCP server supports dynamic toolset selection and visibility filtering at runtime. This allows you to restrict the set of exposed tools in order to prevent blowing up the LLM's context window.

You can configure tool filtering via multiple input channels:

  • CLI Arguments: Pass --tools or --toolsets (or their disabled counterparts --disabled-tools and --disabled-toolsets) during startup.
  • Environment Variables: Define standard environment variables:
  • MCP_ENABLED_TOOLS / MCP_DISABLED_TOOLS
  • MCP_ENABLED_TAGS / MCP_DISABLED_TAGS
  • HTTP SSE Request Headers: Pass custom headers during transport initialization:
  • x-mcp-enabled-tools / x-mcp-disabled-tools
  • x-mcp-enabled-tags / x-mcp-disabled-tags
  • HTTP SSE Request Query Parameters: Append query parameters directly to your transport connection URL:
  • ?tools=tool1,tool2
  • ?tags=tag1

When query strings or parameters are supplied, an LLM-free Knowledge Graph resolution layer (using DynamicToolOrchestrator) matches query intents against known tool tags, names, or descriptions, with safe fallback and automated 24-hour background cache refreshing.


MCP Configuration Examples

> Install the slim [mcp] extra. All examples install audio-transcriber[mcp] — the > MCP-server extra that pulls only the FastMCP / FastAPI tooling (agent-utilities[mcp]). > It deliberately excludes the heavy agent runtime (pydantic-ai, the epistemic-graph > engine, dspy, llama-index), so uvx / container installs are far smaller. Use the > full [agent] extra only when you need the integrated Pydantic AI agent.

stdio Transport (local IDEs — Cursor, Claude Desktop, VS Code)
{
  "mcpServers": {
    "audio-transcriber-mcp": {
      "command": "uvx",
      "args": [
        "--from",
        "audio-transcriber[mcp]",
        "audio-transcriber-mcp"
      ],
      "env": {
        "MCP_TOOL_MODE": "condensed",
        "AUDIO_PROCESSINGTOOL": "True",
        "MISCTOOL": "True",
        "WHISPER_MODEL": "base"
      }
    }
  }
}
Streamable-HTTP Transport (networked / production)
{
  "mcpServers": {
    "audio-transcriber-mcp": {
      "command": "uvx",
      "args": [
        "--from",
        "audio-transcriber[mcp]",
        "audio-transcriber-mcp",
        "--transport",
        "streamable-http",
        "--port",
        "8000"
      ],
      "env": {
        "TRANSPORT": "streamable-http",
        "HOST": "0.0.0.0",
        "PORT": "8000",
        "MCP_TOOL_MODE": "condensed",
        "AUDIO_PROCESSINGTOOL": "True",
        "MISCTOOL": "True",
        "WHISPER_MODEL": "base"
      }
    }
  }
}

Alternatively, connect to a pre-deployed Streamable-HTTP instance by url:

{
  "mcpServers": {
    "audio-transcriber-mcp": {
      "url": "http://localhost:8000/audio-transcriber-mcp/mcp"
    }
  }
}

Deploying the Streamable-HTTP server via Docker:

docker run -d \
  --name audio-transcriber-mcp-mcp \
  -p 8000:8000 \
  -e TRANSPORT=streamable-http \
  -e HOST=0.0.0.0 \
  -e PORT=8000 \
  -e MCP_TOOL_MODE=condensed \
  -e AUDIO_PROCESSINGTOOL=True \
  -e MISCTOOL=True \
  -e WHISPER_MODEL=base \
  knucklessg1/audio-transcriber:mcp

Auto-generated from the code-read env surface (MCP_TOOL_MODE + package vars) — do not edit.

Additional Deployment Options

audio-transcriber can also run as a local container (Docker / Podman / uv) or be consumed from a remote deployment. The Deployment guide has full, copy-paste mcp_config.json for all four transports — stdio, streamable-http, local container / uv, and remote URL:

  • Local container / uv — launch the server from mcp_config.json via uvx,

docker run, or podman run, or point at a local streamable-http container by url.

  • Remote URL — connect to a server deployed behind Caddy at

http://audio-transcriber-mcp.arpa/mcp using the "url" key.

Agent

This repository features a fully integrated Pydantic AI Graph Agent. It communicates over the Agent Control Protocol (ACP) and interacts seamlessly with the Agent Web UI (AG-UI) and Terminal interface.

Running the Agent CLI

To start the interactive command-line agent:

# Configure transcription (optional)
export WHISPER_MODEL="base"
export TRANSCRIBE_DIRECTORY="/path/to/transcribe_directory"

# Run the agent server
audio-transcriber-agent --provider openai --model-id gpt-4o

Docker Compose Orchestration

The following docker/agent.compose.yml configures the Agent, Web UI, and Terminal Interface together:

version: '3.8'

services:
  audio-transcriber-mcp:
    image: knucklessg1/audio-transcriber:mcp
    container_name: audio-transcriber-mcp
    hostname: audio-transcriber-mcp
    restart: always
    env_file:
      - ../.env
    environment:
      - PYTHONUNBUFFERED=1
      - HOST=0.0.0.0
      - PORT=8000
      - TRANSPORT=streamable-http
    ports:
      - "8000:8000"
    healthcheck:
      test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 10s
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"

  audio-transcriber-agent:
    image: knucklessg1/audio-transcriber:latest
    container_name: audio-transcriber-agent
    hostname: audio-transcriber-agent
    restart: always
    depends_on:
      - audio-transcriber-mcp
    env_file:
      - ../.env
    command: [ "audio-transcriber-agent" ]
    environment:
      - PYTHONUNBUFFERED=1
      - HOST=0.0.0.0
      - PORT=9014
      - MCP_URL=http://audio-transcriber-mcp:8000/mcp
      - PROVIDER=${PROVIDER:-openai}
      - MODEL_ID=${MODEL_ID:-gpt-4o}
      - ENABLE_WEB_UI=True
      - ENABLE_OTEL=True
    ports:
      - "9014:9014"
    healthcheck:
      test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:9014/health')"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 10s
    logging:
      driver: json-file
      options:
        max-size: "10m"
        max-file: "3"

Detailed graph node architecture explanations, custom skill configurations, and agentic trace guides are available in [docs/agent.md](docs/agent.md).


Security & Governance

Built directly upon the enterprise-ready agent-utilities core, standard security parameters are fully supported:

Access Control & Policy Enforcement

  • Eunomia Policies: Fine-grained, policy-driven tool authorization. Supports none, local embedded (mcp_policies.json), or centralized remote modes.
  • OIDC Token Delegation: Compliant with RFC 8693 token exchange for flowing authenticating user credentials from Web UI / ACP → Agent → MCP.
  • Scoped Credentials: Execution context runs restricted to the specific caller identity.

Runtime Security Grid

| Feature | Functionality | Enablement | |---------|---------------|------------| | Tool Guard | Sensitivity inspection with human-in-the-loop validation | Enabled by default | | Prompt Injection Defense | Input scanning, repetition monitoring, and recursive loop blocks | Enabled by default | | Context Safety Guard | Stuck-loop detectors and contextual overflow preemptive alerts | Enabled by default |


Environment Variables

Package environment variables

| Variable | Example | Description | |----------|---------|-------------| | HOST | 0.0.0.0 | | | PORT | 8000 | | | TRANSPORT | stdio | options: stdio, streamable-http, sse | | ENABLE_OTEL | True | | | OTEL_EXPORTER_OTLP_ENDPOINT | http://localhost:8080/api/public/otel | | | OTEL_EXPORTER_OTLP_PUBLIC_KEY | pk-... | | | OTEL_EXPORTER_OTLP_SECRET_KEY | sk-... | | | OTEL_EXPORTER_OTLP_PROTOCOL | http/protobuf | | | EUNOMIA_TYPE | none | options: none, embedded, remote | | EUNOMIA_POLICY_FILE | mcp_policies.json | | | EUNOMIA_REMOTE_URL | http://eunomia-server:8000 | | | TRANSCRIBE_DIRECTORY | /path/to/transcribe_directory | Directory where transcripts are written (defaults to the data dir under audio-transcriber) | | MISCTOOL | True | | | AUDIO_PROCESSINGTOOL | True | | | WHISPER_MODEL | base | Standard OpenAI Whisper model to use for local transcription (e.g., base, tiny, small) |

Inherited agent-utilities variables (apply to every connector)

| Variable | Example | Description | |----------|---------|-------------| | MCP_TOOL_MODE | condensed | Tool surface: condensed | verbose | both | | MCP_ENABLED_TOOLS | — | Comma-separated tool allow-list | | MCP_DISABLED_TOOLS | — | Comma-separated tool deny-list | | MCP_ENABLED_TAGS | — | Comma-separated tag allow-list | | MCP_DISABLED_TAGS | — | Comma-separated tag deny-list | | MCP_CLIENT_AUTH | — | Outbound MCP auth (oidc-client-credentials for fleet calls) | | OIDC_CLIENT_ID | — | OIDC client id (service-account auth) | | OIDC_CLIENT_SECRET | — | OIDC client secret (service-account auth) | | DEBUG | False | Verbose logging | | PYTHONUNBUFFERED | 1 | Unbuffered stdout (recommended in containers) | | MCP_URL | http://localhost:8000/mcp | URL of the MCP server the agent connects to | | PROVIDER | openai | LLM provider for the agent | | MODEL_ID | gpt-4o | Model id for the agent | | ENABLE_WEB_UI | True | Serve the AG-UI web interface |

15 package + 14 inherited variable(s). Auto-generated from .env.example + the shared agent-utilities set — do not edit.

Every variable the server reads, grouped by purpose.

Transcription

| Variable | Description | Default | |----------|-------------|---------| | WHISPER_MODEL | Local OpenAI Whisper model (e.g. base, tiny, small) | base | | TRANSCRIBE_DIRECTORY | Directory where transcripts are written | data dir |

MCP server / transport

| Variable | Description | Default | |----------|-------------|---------| | TRANSPORT | stdio, streamable-http, or sse | stdio | | HOST | Bind host (HTTP transports) | 0.0.0.0 | | PORT | Bind port (HTTP transports) | 8000 | | MCP_TOOL_MODE | Tool surface: condensed, verbose, or both | condensed | | MCP_ENABLED_TOOLS / MCP_DISABLED_TOOLS | Comma-separated tool allow/deny list | — | | MCP_ENABLED_TAGS / MCP_DISABLED_TAGS | Comma-separated tag allow/deny list | — | | DEBUG | Verbose logging | False | | PYTHONUNBUFFERED | Unbuffered stdout (recommended in containers) | 1 |

Tool toggles

Each action-routed tool can be disabled individually via its toggle env var (set to false). See the [Available MCP Tools](#available-mcp-tools) table above for the authoritative names.

| Variable | Description | Default | |----------|-------------|---------| | MISCTOOL | Toggle the miscellaneous / health-check tool | True | | AUDIO_PROCESSINGTOOL | Toggle the audio-processing (transcription) tool | True |

Telemetry & governance

| Variable | Description | Default | |----------|-------------|---------| | ENABLE_OTEL | Enable OpenTelemetry export | True | | OTEL_EXPORTER_OTLP_ENDPOINT | OTLP collector endpoint | — | | OTEL_EXPORTER_OTLP_PUBLIC_KEY / OTEL_EXPORTER_OTLP_SECRET_KEY | OTLP auth keys | — | | OTEL_EXPORTER_OTLP_PROTOCOL | OTLP protocol (e.g. http/protobuf) | — | | EUNOMIA_TYPE | Authorization mode: none, embedded, remote | none | | EUNOMIA_POLICY_FILE | Embedded policy file | mcp_policies.json | | EUNOMIA_REMOTE_URL | Remote Eunomia server URL | — |

Agent CLI (full [agent] runtime only)

| Variable | Description | Default | |----------|-------------|---------| | MCP_URL | URL of the MCP server the agent connects to | http://localhost:8000/mcp | | PROVIDER | LLM provider (e.g. openai) | openai | | MODEL_ID | Model id (e.g. gpt-4o) | gpt-4o | | ENABLE_WEB_UI | Serve the AG-UI web interface | True |

See [.env.example](.env.example) for a copy-paste starting point.


Installation

Pick the extra that matches what you want to run:

| Extra | Installs | Use when | |-------|----------|----------| | audio-transcriber[mcp] | Slim MCP server only (agent-utilities[mcp] — FastMCP/FastAPI) | You only run the MCP server (smallest install / image) | | audio-transcriber[agent] | Full agent runtime (agent-utilities[agent,logfire] — Pydantic AI + the epistemic-grap

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.