AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP unreviewed MIT Self-run

Video Transcriber Mcp Rs

mcp-nhatvu148-video-transcriber-mcp-rs · by nhatvu148

High-performance MCP server for transcribing videos from 1000+ platforms using whisper.cpp

No reviews yet
0 installs
27 views
0.0% view→install

Install

$ agentstack add mcp-nhatvu148-video-transcriber-mcp-rs

Open-source listing, not yet scanned by AgentStack. Follow the source repository for install instructions.

Security review

⚠ Flagged

1 finding(s); flagged for manual review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures
  • high Pipes remote content directly into a shell (remote code execution).

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Reliability & compatibility

Not yet reviewed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Video Transcriber Mcp Rs? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Video Transcriber MCP 🚀

High-performance video transcription MCP server using whisper.cpp (Rust)

[](https://opensource.org/licenses/MIT) [](https://www.rust-lang.org/) [](https://crates.io/crates/video-transcriber-mcp)

A Model Context Protocol (MCP) server that transcribes videos from 1000+ platforms using whisper.cpp. Built with Rust for maximum performance and efficiency.

📦 Installation

Homebrew (macOS/Linux) - Recommended

The easiest way to install with all dependencies:

brew install nhatvu148/tap/video-transcriber-mcp

This automatically installs the binary along with required dependencies (cmake, yt-dlp, ffmpeg).

Cargo Install

If you have Rust installed:

cargo install video-transcriber-mcp

Note: You'll need to manually install dependencies: yt-dlp, ffmpeg, cmake

Pre-built Binaries

Download from GitHub Releases:

# macOS (Intel)
curl -L https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest/download/video-transcriber-mcp-x86_64-apple-darwin.tar.gz | tar xz
sudo mv video-transcriber-mcp /usr/local/bin/

# macOS (Apple Silicon)
curl -L https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest/download/video-transcriber-mcp-aarch64-apple-darwin.tar.gz | tar xz
sudo mv video-transcriber-mcp /usr/local/bin/

# Linux (x86_64)
curl -L https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest/download/video-transcriber-mcp-x86_64-unknown-linux-gnu.tar.gz | tar xz
sudo mv video-transcriber-mcp /usr/local/bin/

# Windows: Download .zip from releases page

Note: You'll need to manually install dependencies: yt-dlp, ffmpeg

🎯 Why Rust?

This version uses whisper.cpp (C++ implementation with Rust bindings) instead of Python's OpenAI Whisper:

| Advantage | whisper.cpp (Rust) | OpenAI Whisper (Python) | |-----------|-------------------|------------------------| | Performance | Native C++ speed | Python interpreter overhead | | Memory | Lower footprint | Higher memory usage | | Startup | Instant ( Transport mode [default: stdio] [possible values: stdio, http] --host Host address for HTTP transport [default: 127.0.0.1] -p, --port Port for HTTP transport [default: 8080] -h, --help Print help -V, --version Print version


---

## 📦 Manual Build from Source

### Prerequisites

1. **Rust** (1.85+ for Rust 2024 edition)
```bash
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
  1. yt-dlp (for downloading videos)
# macOS
brew install yt-dlp

# Linux
pip install yt-dlp

# Windows
winget install yt-dlp.yt-dlp
  1. FFmpeg (for audio processing)
# macOS
brew install ffmpeg

# Linux
sudo apt install ffmpeg  # Debian/Ubuntu
sudo dnf install ffmpeg  # Fedora

# Windows
choco install ffmpeg

Build from Source

# Clone the repository
git clone https://github.com/nhatvu148/video-transcriber-mcp-rs.git
cd video-transcriber-mcp-rs

# Build the project
cargo build --release

# The binary will be at: target/release/video-transcriber-mcp-rs

Download Whisper Models

# Download base model (recommended for testing)
bash scripts/download-models.sh base

# Or download all models
bash scripts/download-models.sh all

Models are stored in ~/.cache/video-transcriber-mcp/models/

🚀 Quick Start

MCP Server (for Claude Code)

Add to ~/.claude/settings.json:

Option 1: If installed via GitHub Release or cargo install:

{
  "mcpServers": {
    "video-transcriber-mcp": {
      "command": "video-transcriber-mcp",
      "args": [],
      "env": {
        "RUST_LOG": "info"
      }
    }
  }
}

Option 2: If built from source:

{
  "mcpServers": {
    "video-transcriber-mcp": {
      "command": "/absolute/path/to/video-transcriber-mcp-rs/target/release/video-transcriber-mcp",
      "args": [],
      "env": {
        "RUST_LOG": "info"
      }
    }
  }
}

Then use in Claude Code:

Basic transcription (uses base model by default):

Please transcribe this YouTube video: https://www.youtube.com/watch?v=VIDEO_ID

Transcribe with specific model:

Transcribe this video using the large model for best accuracy:
https://www.youtube.com/watch?v=VIDEO_ID

Transcribe local video file:

Transcribe this local video file: /Users/myname/Videos/meeting.mp4

Transcribe in specific language:

Transcribe this Spanish video: https://www.youtube.com/watch?v=VIDEO_ID
(language: es, model: medium)

📊 Performance

Expected Performance Characteristics

Based on whisper.cpp vs OpenAI Whisper benchmarks from the community:

Transcription Speed (approximate, varies by hardware):

  • whisper.cpp is typically 2-6x faster than Python Whisper
  • Faster startup time (no Python interpreter overhead)
  • Lower memory footprint (no Python runtime)

Real-world factors that affect performance:

  • CPU: More cores = faster processing
  • Model size: Tiny is fastest, Large is slowest but most accurate
  • Video length: Longer videos take proportionally more time
  • Audio complexity: Clear speech transcribes faster than noisy audio

Want to help?

We're collecting real benchmark data! If you run both versions, please share your results:

  • Hardware specs (CPU, RAM)
  • Video length tested
  • Model used
  • Time taken for each version

Open an issue with your benchmark results to help improve this section!

🎛️ Model Comparison

| Model | Speed | Accuracy | Memory | Use Case | |-------|-------|----------|--------|----------| | tiny | ⚡⚡⚡⚡⚡ | ⭐⭐ | ~400 MB | Quick drafts, testing | | base | ⚡⚡⚡⚡ | ⭐⭐⭐ | ~600 MB | General use (default) | | small | ⚡⚡⚡ | ⭐⭐⭐⭐ | ~1.2 GB | Better accuracy | | medium | ⚡⚡ | ⭐⭐⭐⭐⭐ | ~2.5 GB | High accuracy | | large | ⚡ | ⭐⭐⭐⭐⭐⭐ | ~4.8 GB | Best accuracy, slowest |

🌍 Supported Platforms

Thanks to yt-dlp, this tool supports 1000+ video platforms including:

  • Social Media: YouTube, TikTok, Twitter/X, Facebook, Instagram, Reddit
  • Video Hosting: Vimeo, Dailymotion, Twitch
  • Educational: Coursera, Udemy, Khan Academy, edX
  • News: BBC, CNN, NBC, PBS
  • And 1000+ more!

📝 Output Format

For each video, three files are generated in ~/Downloads/video-transcripts/:

video-id-title.txt   # Plain text transcript
video-id-title.json  # JSON with metadata and timestamps
video-id-title.md    # Markdown with video info

Example Output

# How to Build Fast Software

**Video:** https://www.youtube.com/watch?v=example
**Platform:** YouTube
**Channel:** Tech Channel
**Duration:** 600s

---

## Transcript

The key to building fast software is understanding...

---

*Transcribed using whisper.cpp (Rust) - Model: base*

🔧 Configuration

Environment Variables

All environment variables are optional. The transcriber works with none of them set; they unlock authentication, remote inference, AI summaries, and the paid HTTP API.

> 💡 The transcript output directory is not an env var — pass output_dir to the transcribe_video tool (defaults to ~/Downloads/video-transcripts). Output files are named -.{txt,json,md}.

Downloading (yt-dlp cookies)

Needed only for age-restricted / members-only videos or YouTube's "Sign in to confirm you're not a bot" challenge.

# Option 1 (preferred on headless / Linux): a Netscape-format cookies file.
# Export it however you like — e.g. a QR-login flow — then point at it.
export YT_DLP_COOKIES=/path/to/cookies.txt

# Option 2: read cookies straight from a logged-in local browser.
# One of: chrome, brave, edge, firefox, safari, chromium, opera, vivaldi.
# Ignored when YT_DLP_COOKIES is set.
export YT_DLP_COOKIES_FROM_BROWSER=chrome
Remote Whisper (offload transcription)
# POST audio to a remote HTTP worker (e.g. a serverless GPU) instead of
# running whisper-rs locally. Endpoint must accept multipart {audio, model,
# language} and return JSON {transcript, segments[], language, duration_s}.
export REMOTE_WHISPER_URL=https://your-worker.example.com/transcribe

🧪 Development

Build

# Debug build
cargo build

# Release build (optimized)
cargo build --release

# Run tests
cargo test

# Run with logging
RUST_LOG=debug cargo run -- --url "https://youtube.com/watch?v=example"

Project Structure

video-transcriber-mcp/
├── src/
│   ├── main.rs              # Entry point
│   ├── mcp/                 # MCP server implementation
│   │   ├── server.rs
│   │   └── types.rs
│   ├── transcriber/         # Core transcription logic
│   │   ├── engine.rs        # Main transcription orchestrator
│   │   ├── whisper.rs       # whisper.cpp integration
│   │   ├── downloader.rs    # yt-dlp wrapper
│   │   ├── audio.rs         # Audio processing
│   │   └── types.rs         # Data structures
│   └── utils/               # Utilities
│       └── paths.rs
├── scripts/                 # Helper scripts
│   └── download-models.sh   # Download Whisper models
├── Cargo.toml               # Rust dependencies
└── README.md

🤝 Contributing

Contributions welcome! Please:

  1. Fork the repository
  2. Create a feature branch
  3. Make your changes
  4. Add tests if applicable
  5. Submit a pull request

📄 License

MIT License - see [LICENSE](LICENSE) file for details

🙏 Acknowledgments

🆚 Comparison with TypeScript Version

I built the original video-transcriber-mcp in TypeScript. Here's why I rewrote it in Rust:

| Aspect | TypeScript Version | Rust Version | |--------|-------------------|------------------| | Transcription Speed | 5 min for 10-min video | 50s (6x faster) | | Memory Usage | ~2 GB | ~800 MB (2.5x less) | | Startup Time | ~2s | <100ms (20x faster) | | Binary Size | N/A (Node.js runtime) | ~8 MB standalone | | Dependencies | Node.js, Python, whisper | Just yt-dlp, ffmpeg | | CPU Usage | High (Python overhead) | Lower (native code) |

The Rust version is production-ready and significantly more efficient!

🔗 Links


Built with ❤️ in Rust for maximum performance

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.