Install
$ agentstack add skill-serejaris-kimi-skills-audio-generation ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Audio Generation
Use this skill to generate audio. There are two distinct flows — pick the one that matches the user's intent:
- Generate speech (text-to-speech): the user wants spoken audio of some
text. Use the speech flow.
- Generate sound effects: the user wants a sound effect / ambience / SFX
described in words. Use the sound-effects flow.
Setup
Before the first use, ensure the agent-gw Python SDK (version 0.2.6 or newer) is installed. This checks the current environment and installs or upgrades it only when needed:
python3 scripts/audio_generation_tool.py ensure-deps
The SDK needs an API key from api_key=..., KIMI_API_KEY, or ~/.kimi/agent-gw.json.
Choosing the flow
- If the user wants their text read aloud / a voiceover / TTS → speech.
- If the user wants a **sound effect, ambience, music bed, or SFX described in
words → sound-effects**.
Then build the parameters for that flow, run the matching command, and on success surface the saved mp3 to the user. On failure, explain the error from the script; do not invent audio or a local path.
Flow A — Generate speech (text-to-speech)
Parameters:
text(required): the text to convert to speech.voice_id(required): one of the supported voices below. Default is
05Cdh2gw2NMzDvykn1nm.
output(required): local output path ending in.mp3.
Supported voice IDs:
05Cdh2gw2NMzDvykn1nm: calm middle-aged Mandarin male (default)Q63G7WZ5riIGbK8KmqO9: energetic young Mandarin maleNLl76XZRVj1RVeXptX3h: warm Mandarin femaleAt6gj9vUVdJhTriBsuxE: cheerful Mandarin female
Best practices: use punctuation and formatting for natural speech, and break long texts into smaller segments for better quality.
python3 scripts/audio_generation_tool.py speech \
--text "你好,欢迎使用 Kimi。" \
--voice-id "05Cdh2gw2NMzDvykn1nm" \
--output "/path/to/output.mp3"
This sends {"text", "voice_id"} to the gateway generate_speech API.
Flow B — Generate sound effects
Parameters:
description(required): a detailed description of the sound effect. It
MUST be in English — never use another language.
duration(required): duration in seconds, range0.5-22.output(required): local output path ending in.mp3.
Duration guidance: short (0.5-3s) for UI sounds/notifications, medium (3-10s) for ambient loops or action sequences, long (10-22s) for background music or extended ambience.
python3 scripts/audio_generation_tool.py sound-effects \
--description "Gentle rain falling on leaves with distant thunder" \
--duration 8 \
--output "/path/to/output.mp3"
This sends {"description", "duration_seconds"} to the gateway generate_sound_effects API.
After generation
Both flows read media.url / media.mime_type from the response and download the audio to your output path with curl (allowing up to 5 minutes). The script then prints the saved mp3 path. Surface that audio to the user (e.g. via the readFile tool on the path). Playing or reading the audio to present it is the model's job, not this plugin's work.
Script summary
The script:
speechvalidatesvoice_idagainst the supported list and sends
{"text", "voice_id"} to generate_speech
sound-effectsvalidates the duration is within0.5-22and sends
{"description", "duration_seconds"} to generate_sound_effects
- reads the generated
media.urlandmedia.mime_typefrom the response - downloads the audio to the
--outputpath withcurl(up to 5 minutes),
naming the file as .mp3
- prints the saved path and a reminder to surface it to the user
Response shape for both APIs (resp.json()):
{
"media": {
"url": str, # public URL of the generated audio
"mime_type": str, # e.g. "audio/mpeg"
}
}
> This skill uses the agent-gw Python SDK: > client.tools.generate_speech(text, voice_id=...) and > client.tools.generate_sound_effects(description, duration_seconds=...), with > the response shape above.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: serejaris
- Source: serejaris/kimi-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.