AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Text To Speech

skill-claude-world-claude-agent-text-to-speech · by claude-world

>

No reviews yet
0 installs
9 views
0.0% view→install

Install

$ agentstack add skill-claude-world-claude-agent-text-to-speech

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-claude-world-claude-agent-text-to-speech)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Text To Speech? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Text to Speech

Convert text to audio. Three engine options from instant offline to high-quality cloud.

Engines Overview

| Engine | Quality | Speed | Cost | Install | |--------|---------|-------|------|---------| | say (macOS) | Decent | Instant | Free | Built-in | | sherpa-onnx-tts | Good | Fast | Free | pip install | | ElevenLabs API | Excellent | ~2s | Paid | API key |

Default: use say for quick playback, save to file with -o.

Engine 1: macOS say (Built-in, No Install)

# Speak immediately
say "Hello, world!"

# Save to audio file
say "Your message here" -o output.aiff

# Convert to MP3 (requires ffmpeg)
say "Your message here" -o /tmp/tts.aiff && ffmpeg -y -i /tmp/tts.aiff output.mp3

# List available voices
say -v ?

# Use specific voice
say -v Samantha "Hello from Samantha"
say -v Daniel "Hello from Daniel"
say -v Mei-Jia "你好,這是中文語音"
say -v Kyoko "こんにちは、日本語です"

# Control speed (default 200 wpm)
say -r 150 "Slower speech"
say -r 250 "Faster speech"

# Read from file
say -f input.txt -o output.aiff

Popular macOS Voices

| Language | Voice | Note | |----------|-------|------| | English | Samantha | Default US | | English | Daniel | UK accent | | Chinese | Mei-Jia | Traditional Chinese | | Chinese | Ting-Ting | Simplified Chinese | | Japanese | Kyoko | Japanese |

Engine 2: sherpa-onnx-tts (Local Neural TTS)

# Install
pip install sherpa-onnx

# Download a model (example: VITS English)
# See: https://k2-fsa.github.io/sherpa/onnx/tts/index.html
sherpa-onnx-tts --help

# Generate audio
sherpa-onnx-tts \
  --vits-model=./model.onnx \
  --vits-lexicon=./lexicon.txt \
  --vits-tokens=./tokens.txt \
  --output-filename=output.wav \
  "Text to convert to speech"

Engine 3: ElevenLabs API (Highest Quality)

# Requires: export ELEVENLABS_API_KEY=your_key_here

# List voices
curl -s https://api.elevenlabs.io/v1/voices \
  -H "xi-api-key: $ELEVENLABS_API_KEY" | jq '.voices[].name'

# Generate audio (Rachel voice)
curl -s -X POST \
  "https://api.elevenlabs.io/v1/text-to-speech/21m00Tcm4TlvDq8ikWAM" \
  -H "xi-api-key: $ELEVENLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "Your text here", "model_id": "eleven_monolingual_v1"}' \
  -o output.mp3

Usage Examples

"Read this out loud: Good morning" (immediate playback)

say "Good morning"

"Save 'Meeting starts in 5 minutes' to audio"

say "Meeting starts in 5 minutes" -o ~/Desktop/reminder.aiff

"Generate MP3 from this text"

say "Your text here" -o /tmp/tts_out.aiff
ffmpeg -y -i /tmp/tts_out.aiff ~/Desktop/output.mp3 2>/dev/null
echo "Saved: ~/Desktop/output.mp3"

"What voices are available?"

say -v ? | head -20

"Read this in Chinese"

say -v Mei-Jia "你好,歡迎使用語音服務"

Rules

  • say is always available on macOS — use it as default, no install check needed
  • For file output: default save to ~/Desktop/tts_output_.aiff
  • Convert to MP3 only if user requests it and ffmpeg is available
  • If user wants higher quality, suggest sherpa-onnx or ElevenLabs
  • For ElevenLabs, check ELEVENLABSAPIKEY is set before attempting API call
  • Never hardcode API keys — always read from environment variable
  • Long text (>500 words): warn user it may take a moment, especially for sherpa/ElevenLabs
  • For non-English, default to the appropriate system voice (Mei-Jia for ZH, Kyoko for JA)

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.