Install
$ agentstack add skill-godot-fun-godot-framework-ai-text-to-speech ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
AI Text-to-Speech (IndexTTS2)
Clone a speaker from a reference audio, then synthesize speech from text with IndexTTS2.
Rules
When this skill applies, read and follow [skill-dependency-manager](../../rules/skill-dependency-manager.md) — run scripts as documented, install missing tools into .dependency/.
- Run
tts.pythrough theindex-ttsmanifest entry (.dependency/index-tts/.venv/). Never use hostpython,py,python3, or any interpreter outside.dependency/. - Do not hand-write IndexTTS Python snippets or
uv run webui.pyfor synthesis — use the bundled script. - IndexTTS requires
uvfor install (pip/condaare unsupported upstream). Python must be **>=3.10,/tts/speech.wav):
.dependency/index-tts/.venv/Scripts/python.exe .cursor/skills/ai-text-to-speech/scripts/tts.py \
--voice audio/voice/ref.wav \
--text "你好,欢迎来到这个世界。"
# → audio/voice/tts/speech.wav
Explicit output path:
.dependency/index-tts/.venv/Scripts/python.exe .cursor/skills/ai-text-to-speech/scripts/tts.py \
--voice audio/voice/ref.wav \
--text "Hello, this is a test." \
--output audio/voice/tts/hello.wav
Long script from a UTF-8 text file:
.dependency/index-tts/.venv/Scripts/python.exe .cursor/skills/ai-text-to-speech/scripts/tts.py \
--voice audio/voice/ref.wav \
--text-file script/lines/intro.txt \
--output audio/voice/tts/intro.wav
FP16 (faster, less VRAM):
.dependency/index-tts/.venv/Scripts/python.exe .cursor/skills/ai-text-to-speech/scripts/tts.py \
--voice audio/voice/ref.wav \
--text "测试半精度推理。" \
--fp16
Emotion control (optional)
| Mode | Flags | Notes | |------|-------|-------| | Emotion reference audio | --emotion-audio path.wav | Separate clip for emotion; timbre still from --voice | | Emotion weight | --emotion-weight 0.6 | Maps to emo_alpha (0.0–1.0, default 1.0) | | Emotion from text | --emotion-from-text | Infer emotion from synthesis text; prefer --emotion-weight ≈ 0.6 | | Emotion description | --emotion-text "..." | Natural-language emotion; implies text emotion mode | | Emotion vector | --emotion-vector 0,0,0.8,0,0,0,0,0 | 8 floats: happy, angry, sad, afraid, disgusted, melancholic, surprised, calm |
# Emotion reference audio
.dependency/index-tts/.venv/Scripts/python.exe .cursor/skills/ai-text-to-speech/scripts/tts.py \
--voice audio/voice/ref.wav \
--emotion-audio audio/voice/emo_sad.wav \
--emotion-weight 0.9 \
--text "酒楼丧尽天良,开始借机竞拍房间。" \
--output audio/voice/tts/sad_line.wav
# Emotion description text
.dependency/index-tts/.venv/Scripts/python.exe .cursor/skills/ai-text-to-speech/scripts/tts.py \
--voice audio/voice/ref.wav \
--emotion-text "害怕、紧张" \
--emotion-weight 0.6 \
--text "快躲起来!是他要来了!" \
--output audio/voice/tts/afraid_line.wav
Do not combine --emotion-audio, --emotion-vector, and --emotion-text / --emotion-from-text in conflicting ways — pick one emotion source.
Defaults
| Option | Default | Notes | |--------|---------|-------| | Output | /tts/speech.wav | Auto-create tts/; use --output to override | | Model | .dependency/index-tts/checkpoints | IndexTTS-2 | | --fp16 | off | Enable on GPU when VRAM is tight | | --emotion-weight | 1.0 | Lower (~0.6) for text emotion modes | | Overwrite | off | Pass --force to replace an existing output |
Agent workflow
- Confirm inputs — need a clear reference voice WAV/MP3 and the text (or
--text-file). Ask if either is missing. - Use the user's real paths — do not copy voice files into the repo unless asked.
- Trial first — synthesize one short line, play/inspect before long scripts.
- Reference audio tips — clean, single-speaker, little noise; a few seconds of clear speech works best.
- Missing install — follow Setup; register
index-ttsinmanifest.json; retry the same command. - GPU — prefer CUDA +
--fp16for speed; CPU is acceptable for short tests only. - Revert — delete files under
tts/; sources are never modified.
Troubleshooting
| Issue | Fix | |-------|-----| | index-tts not populated | Clone + uv sync + download checkpoints; update manifest | | checkpoints/config.yaml missing | Re-run hf download IndexTeam/IndexTTS-2 --local-dir=checkpoints | | CUDA / torch errors | Install CUDA 12.8+; or run on CPU (slow) | | OOM / VRAM | Pass --fp16; shorten text; close other GPU apps | | Slow HuggingFace | Set HF_ENDPOINT=https://hf-mirror.com; or use ModelScope | | uv sync / DeepSpeed fail on Windows | Use uv sync --extra webui without deepspeed | | Unnatural emotion | Lower --emotion-weight to ~0.6; try a clearer --emotion-audio |
Related
- Upstream: https://github.com/index-tts/index-tts
- Models: IndexTTS-2 (HuggingFace)
- Post-process loudness / format: [audio-loudness-normalization](../audio-loudness-normalization/SKILL.md), [audio-to-ogg](../audio-to-ogg/SKILL.md), [audio-to-wav](../audio-to-wav/SKILL.md)
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: godot-fun
- Source: godot-fun/godot-framework
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.