AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Media Generation

skill-onimusya-media-gen-media-generation · by onimusya

Generate and edit images, create videos, synthesize speech, transcribe and translate audio using multiple AI providers (OpenAI, Google, ElevenLabs, Deepgram, Fal, Luma, Replicate, Stability, Runway, OpenRouter, Edge TTS). Use when the user asks to create media assets, generate pictures, make videos, produce voiceovers, transcribe recordings, or work with any visual/audio content.

— No reviews yet
0 installs
25 views
0.0% view→install

Install

$ agentstack add skill-onimusya-media-gen-media-generation

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ✓ Network access No
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ● Environment & secrets Used
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-onimusya-media-gen-media-generation)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● 2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Media Generation? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Media Generation

Run the CLI at ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs with Node.js. Always use --json for parseable output.

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs  [options] --json

Commands

| Command | Purpose | |---------|---------| | image generate | Generate images from text prompts | | image edit | Edit existing images with a prompt | | video generate | Generate video from text (async) | | video image-to-video | Animate an image into video | | video extend | Extend an existing video | | voice tts | Text to speech synthesis | | voice clone | Clone a voice from audio samples | | voice isolate | Isolate voice from background audio | | audio transcribe | Transcribe audio to text | | audio translate | Translate audio to another language | | providers list | List all providers with capabilities, models, and voices | | providers list --configured | List only configured providers | | providers list --capability | Filter by capability | | providers models | List all models across all providers | | providers models --provider | List models/voices for a specific provider | | providers models --capability | Filter models by capability | | config init | Initialize project-level config | | config init --global | Initialize user-level config at ~/.media-gen/ | | config validate | Check which providers are configured | | job status | Check async job status | | job download | Download completed async job result |

Configuration

Defaults are set via .env (at project root or ~/.media-gen/.env for global). With defaults configured, --provider, --model, and --voice-id are all optional:

MEDIA_GEN_DEFAULT_PROVIDER=openrouter
MEDIA_GEN_DEFAULT_MODEL=openai/gpt-image-2
MEDIA_GEN_VOICE_PROVIDER=edge-tts
MEDIA_GEN_VOICE_MODEL=
MEDIA_GEN_VOICE_ID=en-US-EmmaMultilingualNeural
MEDIA_GEN_VIDEO_PROVIDER=google
MEDIA_GEN_VIDEO_MODEL=veo-3.1-generate-preview
MEDIA_GEN_AUDIO_PROVIDER=deepgram
MEDIA_GEN_AUDIO_MODEL=nova-3
MEDIA_GEN_LOG_LEVEL=error

Examples

Image generation

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs image generate \
  --prompt "A pixel art fantasy arena" \
  --output ./outputs/arena.png \
  --json

Video generation (wait for result)

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs video generate \
  --provider google \
  --model veo-3.1-generate-preview \
  --prompt "A cinematic card pack opening" \
  --duration 8 \
  --output ./outputs/video.mp4 \
  --wait \
  --json

Text to speech (Edge TTS - free, no API key)

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs voice tts \
  --provider edge-tts \
  --voice-id en-US-EmmaMultilingualNeural \
  --text "Hello world" \
  --output ./outputs/voice.mp3 \
  --json

Text to speech (ElevenLabs)

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs voice tts \
  --provider elevenlabs \
  --voice-id JBFqnCBsd6RMkjVDRZzb \
  --text "Welcome to the show" \
  --output ./outputs/george.mp3 \
  --json

Text to speech (Google Gemini TTS)

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs voice tts \
  --provider google \
  --model gemini-3.1-flash-tts-preview \
  --voice-id Kore \
  --text "Say cheerfully: Have a wonderful day!" \
  --output ./outputs/gemini-voice.wav \
  --json

Text to speech (OpenAI)

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs voice tts \
  --provider openai \
  --model gpt-4o-mini-tts \
  --voice-id coral \
  --text "Hello from OpenAI" \
  --output ./outputs/openai-voice.mp3 \
  --json

Text to speech (with defaults set, minimal)

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs voice tts \
  --text "Just provide text when defaults are configured" \
  --output ./outputs/speech.mp3 \
  --json

Transcription

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs audio transcribe \
  --input ./audio/recording.mp3 \
  --output ./outputs/transcript.json \
  --json

List supported providers and models

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs providers list --json

List voices for a TTS provider

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs providers models --provider edge-tts --json
node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs providers models --provider elevenlabs --json
node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs providers models --provider openai --json
node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs providers models --provider google --json

Dry run (validate without calling API)

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs image generate \
  --prompt "test" --dry-run --json

Async Video Jobs

Video generation is asynchronous. Providers return a job ID instead of a file.

Pattern 1: Wait for completion (simple)

Add --wait to block until the video is ready:

node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs video generate \
  --provider google \
  --model veo-3.1-generate-preview \
  --prompt "A cinematic scene" \
  --output ./outputs/video.mp4 \
  --wait \
  --poll-interval 5000 \
  --timeout 300000 \
  --json

Returns the final file path when complete.

Pattern 2: Non-blocking (get job ID, check later)

Without --wait, the CLI returns immediately with a job ID:

# Start generation (returns instantly)
node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs video generate \
  --provider google \
  --model veo-3.1-generate-preview \
  --prompt "A cinematic scene" \
  --json
# Returns: {"ok": true, "jobId": "operations/abc123", "status": "processing"}

# Check status later
node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs job status \
  --provider google \
  --job-id "operations/abc123" \
  --json
# Returns: {"ok": true, "jobId": "...", "status": "completed"}

# Download the result
node ${CLAUDE_SKILL_DIR}/scripts/media-gen.mjs job download \
  --provider google \
  --job-id "operations/abc123" \
  --output ./outputs/video.mp4 \
  --json

Async options

| Option | Default | Description | |--------|---------|-------------| | --wait | false | Block until job completes | | --poll-interval | 5000 | Milliseconds between status checks | | --timeout | 600000 | Max wait time (10 minutes) |

Async providers

Google (Veo), Luma AI, Runway, Fal.ai, and Replicate all use async for video. Image and TTS are always synchronous.

Response format

Success:

{"ok": true, "type": "image", "provider": "openai", "model": "gpt-image-2", "outputFile": "./outputs/image.png", "durationMs": 1200}

Error:

{"ok": false, "error": {"code": "PROVIDER_NOT_CONFIGURED", "message": "Missing OPENAI_API_KEY", "suggestion": "Set OPENAI_API_KEY in .env"}}

Rules

  • Use --json for all calls.
  • Use --dry-run before expensive operations when unsure.
  • Never use --overwrite unless the user confirms.
  • Keep outputs inside the project workspace.
  • For video, use --wait only when the user wants the file immediately.
  • Check the ok field in every response before proceeding.
  • On error, show the suggestion field to the user.
  • For TTS, --voice-id is optional when MEDIA_GEN_VOICE_ID is set in .env.
  • Edge TTS is free and requires no API key — prefer it for basic TTS tasks.

Provider Capabilities

| Provider | Image | Video | TTS | Transcribe | Translate | Clone | Isolate | |----------|-------|-------|-----|------------|-----------|-------|---------| | openai | Yes | | Yes | Yes | Yes | | | | google | Yes | Yes | Yes | | | | | | azure | Yes | | Yes | Yes | Yes | | | | elevenlabs | | | Yes | | | Yes | Yes | | deepgram | | | Yes | Yes | Yes | | | | fal | Yes | Yes | | | | | | | luma | | Yes | | | | | | | replicate | Yes | Yes | | | | | | | stability | Yes | | | | | | | | runway | | Yes | | | | | | | openrouter | Yes | Yes | | | | | | | edge-tts | | | Yes (free) | | | | |

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.