# Elevenlabs Tts

> Generate Chinese-friendly speech from text via ElevenLabs API. Supports one-shot TTS, multi-voice dialogue/podcast generation, v3 timestamped dialogue (for karaoke-synced transcripts), and multi-chunk merging. Use when the user asks to make audio narration, podcasts, voice-overs, interviews, or any 中文/英文 TTS task that calls out 「ElevenLabs」 / 「11labs」 / 「TTS」 / 「配音」 / 「播客」. Read references/v3-pod…

- **Type:** Skill
- **Install:** `agentstack add skill-duoduo25-elevenlabs-tts-skill-elevenlabs-tts-skill`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [DuoDuo25](https://agentstack.voostack.com/s/duoduo25)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [DuoDuo25](https://github.com/DuoDuo25)
- **Source:** https://github.com/DuoDuo25/elevenlabs-tts-skill

## Install

```sh
agentstack add skill-duoduo25-elevenlabs-tts-skill-elevenlabs-tts-skill
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# elevenlabs-tts

A thin Python wrapper around the ElevenLabs TTS HTTP API. **Pure stdlib only** — no `pip install` needed. Outputs MP3 (+ optional character-level timing JSON).

Built and battle-tested on a real 17-minute Chinese interview podcast (5 acts, 18MB audio, character-level karaoke sync). All the踩坑 lessons live in `references/v3-podcast-cookbook.md`.

## When to use

- Single-line narration / voice-over (`scripts/tts.py`)
- Multi-voice dialogue / podcast / interview, simple concat (`scripts/dialogue.py`)
- Multi-voice **with character-level timestamps** for karaoke-synced UI (`scripts/dialogue_timestamps.py`)
- Merge multiple TTS chunks into a single seekable episode (`scripts/merge_chunks.py`)
- Discover what voices the user's account has (`scripts/list_voices.py`)

If the user wants Whisper/transcription, voice cloning, or sound effects — don't use this skill, use a specialized one.

## 🔥 BEFORE WRITING ANY 中文 SCRIPT

**Read `references/v3-podcast-cookbook.md`.** It has the hard-won rules from a real podcast build:

- v3 audio tag patterns that actually move the model
- Speech-speed control via tags (v3 has no `speed` param)
- Punctuation rules to control pacing (`。` vs `，`, kill all `——`)
- TTS literacy rules for Chinese (`25 年` → `二五年`, `B+` → `B 加`, `早九晚十一半` → `早九点上班、晚十一点半下班`)
- HOST vs GUEST tone calibration
- Concurrency limits + chunking
- The MP3 concat `-c copy` bug that silently breaks browser seek
- The CDN Range bug + Blob URL fix for static-asset hosting

Skipping it = burning credits on regenerations.

## Setup (one-time)

The skill reads the API key in this order:

1. `ELEVENLABS_API_KEY` env var
2. `~/.config/elevenlabs/api_key` (single line, no quotes)

```bash
mkdir -p ~/.config/elevenlabs && \
  printf '%s' "" > ~/.config/elevenlabs/api_key && \
  chmod 600 ~/.config/elevenlabs/api_key
```

**Never commit your key.** The provided `.gitignore` excludes common key paths, but the canonical location is outside the repo on purpose.

## Core command · single-line TTS

```bash
python3 scripts/tts.py \
  --voice  \
  --text "你好世界，这是一段测试。" \
  --out /tmp/hello.mp3 \
  [--model eleven_multilingual_v2] \
  [--stability 0.5] [--similarity 0.75] [--style 0] [--speed 1.0] \
  [--format mp3_44100_128]
```

- `--voice` accepts a raw `voice_id` or a shorthand from `references/voices.md`.
- Defaults work for 中文: `--model eleven_multilingual_v2`, `--stability 0.5`, `--similarity 0.75`, `--style 0`.

## Dialogue / podcast command · basic, no timestamps

Input is either `.txt` (`SPEAKER: text` per line) or `.json` (with per-speaker settings).

```bash
python3 scripts/dialogue.py \
  --script /path/to/script.txt \
  --voices "HOST=adam,GUEST=rachel" \
  --out /path/to/podcast.mp3 \
  [--gap-ms 350] [--keep-parts] [--model eleven_multilingual_v2]
```

Output: `podcast.mp3` + `podcast.json` sidecar with per-line duration/timecode (line-level, NOT character-level).

## v3 dialogue with character-level timestamps · for karaoke UI

```bash
python3 scripts/dialogue_timestamps.py \
  /path/to/script.json \
  /path/to/out_basename
```

`script.json` shape (call this the **v3 dialogue body**):

```json
{
  "model_id": "eleven_v3",
  "settings": {"stability": 0.0},
  "language_code": "zh",
  "inputs": [
    {"voice_id": "", "text": "[curious, warmly] "},
    {"voice_id": "", "text": "[matter-of-factly] "}
  ]
}
```

Output:

- `out_basename.mp3` — the audio
- `out_basename.timing.json` — `{alignment, voice_segments, script_lines}` for sync UIs

Pattern + UI tips in cookbook §1 and §7.

## Merge multiple chunks into one seekable episode

Long podcasts must be chunked (~500-800 chars per call) or they time out. After generating each chunk:

```bash
python3 scripts/merge_chunks.py \
  /path/to/episode_out \
  /path/to/chunk_01 \
  /path/to/chunk_02 \
  ...
```

Reads each `.mp3` + `.timing.json`, writes:

- `episode_out.mp3` — concatenated audio (re-encoded with `libmp3lame` so the global Xing/Info VBR header is consistent — `-c copy` silently breaks browser seek; cookbook §6)
- `episode_out.timing.json` — merged alignment with all offsets time-shifted onto the global timeline

## Listing voices

```bash
python3 scripts/list_voices.py [--search "..."] [--category premade|cloned|professional] [--page-size 50]
```

Prints TSV. Use before generating to verify the voice_id is valid + viewable to the user.

## Choosing the model

| Model | When | Notes |
|---|---|---|
| `eleven_multilingual_v2` | **Default for 中文 narration / single-voice / quick batch** | Stable, fast, good 中文 |
| `eleven_v3` | **Multi-voice dialogue podcast with emotion tags + timestamps** | Required for audio tags `[bitter]`/`[quickly]` to work; pair with `settings.stability=0.0` (Creative). See cookbook. |
| `eleven_turbo_v2_5` | When user mentions latency / streaming | Lower quality, faster latency — only when user explicitly wants speed |

## Voice catalog (premade · always available)

See `references/voices.md` for the full catalog of premade voice IDs plus suggested HOST/GUEST pairings for Chinese podcasts. Quick recall:

| Shorthand | voice_id | Best for |
|---|---|---|
| `rachel` | `21m00Tcm4TlvDq8ikWAM` | Calm female narrator |
| `bella` | `EXAVITQu4vr4xnSDxMaL` | Soft female |
| `adam` | `pNInz6obpgDQGcFmaJgB` | Deep male narrator |
| `antoni` | `ErXwobaYiN019PkySvjV` | Warm well-rounded male |
| `josh` | `TxGEqnHWrfWFTfGW9XjX` | Deep American male |
| `arnold` | `VR6AewLTigWG4xSOukaG` | Crisp older male |
| `domi` | `AZnzlk1XvdvUeBnXmlld` | Strong confident female |
| `elli` | `MF3mGyEYCl7XYWbV9V6O` | Emotional female |
| `sam` | `yoZ06aMxZJJ28mfd3POQ` | Raspy young male |
| `charlotte` | `XB0fDUnXU5powFXDhCwa` | Warm British female |

⚠️ **Library / shared voices need a paid plan**. Free tier returns `402 paid_plan_required` even if voice was saved to the account. Only premade / cloned voices work on free.

## Tunable settings · quick guidance

| Setting | Range | Effect |
|---|---|---|
| `stability` | 0.0–1.0 | Lower = more emotional/varied. **Chinese narrative: 0.4–0.55** for v2; **0.0 (Creative)** for v3 with tags |
| `similarity_boost` | 0.0–1.0 | How tightly to match the original voice. **0.7–0.85** sweet spot |
| `style` | 0.0–1.0 | Style exaggeration. **≤ 0.2** for podcast/narration; higher = caricature |
| `speed` | 0.7–1.2 | v2 only — v3 ignores this, control speed via tags instead |
| `use_speaker_boost` | bool | Default `true` |

## Cost awareness

ElevenLabs charges by character. **Before running a large generation, give the user a rough character count.** A typical 20-min podcast is ~6,000-8,000 chars; multilingual_v2 costs ~2× per char vs English. v3 costs same.

⚠️ **Tags count toward character billing.** `[bitter chuckle, quickly]` is 23 chars. Heavy tagging can double script length.

## Hard rules / anti-patterns

- ❌ **Don't use `ffmpeg -c copy` to concat MP3 chunks** — only the first chunk's Xing/Info VBR header gets preserved, breaking browser seek. Use `ffmpeg -c:a libmp3lame -b:a 128k -ar 44100 -ac 1` to re-encode and get a fresh global Xing. `scripts/merge_chunks.py` does this for you.
- ❌ **Don't fire >3 generations in parallel** — EL has a 3-concurrent-request limit; the 4th gets `429 concurrent_limit_exceeded`. Batch 2 + 2.
- ❌ **Don't generate a 7,000-char podcast in one shot before letting the user hear a sample.** Generate the cold open first, get sign-off on voice/tone, then continue.
- ❌ **Don't write the API key into a tracked file** — only into `~/.config/elevenlabs/api_key` or env.
- ❌ **Don't `set -e` past failed TTS calls** — surface ElevenLabs error JSON verbatim.
- ❌ **Don't reach for `eleven_v3` casually** — only when you explicitly want audio-tag emotion control. For plain narration, multilingual_v2 is faster + cheaper.
- ❌ **Don't write arabic years/dates in 中文 scripts** (`25 年` reads as 二十五年, dates get mangled). See cookbook §4.
- ❌ **Don't write `B+` / `B-` / 「早九晚十一半」** — TTS mispronounces. See cookbook §4.
- ❌ **Don't request `audio.currentTime = N` on a CDN that doesn't serve Range requests** — browser can't seek past the buffered region. Fetch as Blob + `URL.createObjectURL` if deploying to a Range-hostile CDN (e.g. Cloudflare Workers Static Assets, EdgeSpark). See cookbook §6.
- ❌ **Don't drive karaoke highlight off `.ontimeupdate`** — it fires only ~4Hz, lag is visible. Use `requestAnimationFrame` to poll `audio.currentTime` at ~60Hz. See cookbook §7.

## References

| File | When to read |
|---|---|
| `references/v3-podcast-cookbook.md` | **Always before writing any 中文 podcast script** — tags, pacing, literacy rules, merge bugs, sync UI tips, real templates |
| `references/voices.md` | Picking a voice + premade ID lookup + suggested HOST/GUEST pairings |

## License

MIT — see `LICENSE`. Use freely, including commercially.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [DuoDuo25](https://github.com/DuoDuo25)
- **Source:** [DuoDuo25/elevenlabs-tts-skill](https://github.com/DuoDuo25/elevenlabs-tts-skill)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-duoduo25-elevenlabs-tts-skill-elevenlabs-tts-skill
- Seller: https://agentstack.voostack.com/s/duoduo25
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
