AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Foundry Voice Live

skill-aiappsgbb-awesome-gbb-foundry-voice-live · by aiappsgbb

>

No reviews yet
0 installs
16 views
0.0% view→install

Install

$ agentstack add skill-aiappsgbb-awesome-gbb-foundry-voice-live

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-aiappsgbb-awesome-gbb-foundry-voice-live)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Foundry Voice Live? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Foundry Voice Live

Build real-time voice agents on Azure AI Foundry using Voice Live — the GA (2026-04-10) server-side voice pipeline that adds semantic VAD, echo cancellation, noise reduction, and Azure Neural HD voices on top of the standard Azure OpenAI Realtime API.

The migration from Realtime to Voice Live is three small code changes; the migration from openai-shim to the native azure-ai-voicelive SDK is one additional step (Rung 4). This skill walks through both ladders, the session config that unlocks Voice Live features, and the four 2026-04-10 GA deltas — proactive turn control, MCP tools mid-turn, OpenTelemetry via diagnostic settings, and auto-truncate governance.

When to Use

Use this skill when the task involves ANY of:

  • Real-time speech-to-speech interaction with a model
  • Migrating from Azure OpenAI Realtime to Voice Live
  • Adding voice to a Foundry hosted or prompt agent
  • Building a voice demo with Gradio / FastRTC / WebRTC
  • Benchmarking realtime voice latency (TTFA, TTFT)
  • Comparing Azure Neural HD voices vs OpenAI voices

Do NOT use for batch STT/TTS (foundry-doc-vision-speech), non-voice agents (foundry-hosted-agents, foundry-prompt-agents), or document extraction.


1 · The Four Rungs

The entire migration from "plain Realtime" to "native Voice Live SDK on a Foundry Agent" is a diff ladder — each rung changes only the connection-setup block.

Rung 1: Azure OpenAI Realtime              ← the "before"
  │
  ▼  diff = 3 small lines (api_version, websocket_base_url, extra_query)
Rung 2: Azure Voice Live                   ← the punchline
  │
  ▼  diff = 1 line (extra_query gains agent-id / -project-name / -access-token)
Rung 3: Voice Live + Foundry Agent         ← the endgame on `openai` SDK
  │
  ▼  swap `openai.AsyncAzureOpenAI` → `azure.ai.voicelive.aio.connect`
Rung 4: native `azure-ai-voicelive` SDK    ← the GA path

Everything else — the audio pipe, transcript fan-out, status events, voice picker, UI — is identical across all four rungs.


2 · Connection Code (Copy-Paste Ready)

Rung 1 — Azure OpenAI Realtime

from contextlib import asynccontextmanager
from openai import AsyncAzureOpenAI

@asynccontextmanager
async def connect_realtime(*, settings, token_provider):
    client = AsyncAzureOpenAI(
        azure_endpoint=settings.azure_endpoint,              # https://.openai.azure.com
        api_version="2025-04-01-preview",                    # Realtime preview
        azure_ad_token_provider=token_provider,
    )
    try:
        async with client.realtime.connect(
            model=settings.azure_deployment_name,
        ) as conn:
            yield conn
    finally:
        await client.close()

Rung 2 — Azure Voice Live (3 changed lines)

@asynccontextmanager
async def connect_voicelive(*, settings, token_provider):
    actual_model = settings.azure_deployment_name
    client = AsyncAzureOpenAI(
        azure_endpoint=settings.azure_endpoint,
        api_version="2026-04-10",                                       # ← GA
        azure_ad_token_provider=token_provider,
        websocket_base_url=settings.azure_voice_live_endpoint,          # ← wss://.../voice-live
    )
    try:
        async with client.realtime.connect(
            model=actual_model,
            extra_query={"model": actual_model},                        # ← &model= not &deployment=
        ) as conn:
            yield conn
    finally:
        await client.close()

Rung 3 — Voice Live + Foundry Agent (extra_query extends)

@asynccontextmanager
async def connect_agent(*, settings, token_provider, agent_token_provider):
    client = AsyncAzureOpenAI(
        azure_endpoint=settings.azure_endpoint,
        api_version="2026-04-10",
        azure_ad_token_provider=token_provider,
        websocket_base_url=settings.azure_voice_live_endpoint,
    )
    try:
        async with client.realtime.connect(
            model=settings.azure_deployment_name,
            extra_query={
                "agent-id":           settings.agent_id,
                "agent-project-name": settings.agent_project_name,
                "agent-access-token": await agent_token_provider(),   # ai.azure.com scope
            },
        ) as conn:
            yield conn
    finally:
        await client.close()

Why 3 lines, not 1

  1. websocket_base_url — the headline change; redirects the WSS

connection to /voice-live on services.ai.azure.com.

  1. api_version — Realtime is still on 2025-04-01-preview

(the openai 2.x SDK emits /openai/realtime; when it adopts the GA /openai/v1/realtime URL this difference collapses). Voice Live is GA — the latest stable version is 2026-04-10 (was 2025-10-01 pre-//build 2026).

  1. extra_query={"model": ...} — the SDK adds &deployment=… to

the WSS URL by default; Voice Live keys off &model=…, so we add it explicitly.

Rung 4 — native azure-ai-voicelive SDK

The azure-ai-voicelive Python SDK (stable 1.2.0; latest preview 1.3.0b1) is the first-party path for Voice Live. It speaks the same wire protocol as Rungs 2–3 (so the event-handling code in §9 still works verbatim), but replaces the openai-shim plumbing with a typed, Voice-Live-native client.

from contextlib import asynccontextmanager

from azure.ai.voicelive.aio import connect       # async client
from azure.ai.voicelive.models import (
    AzureSemanticVad,
    AzureStandardVoice,
    InputAudioFormat,
    Modality,
    OutputAudioFormat,
    RequestSession,
)
from azure.identity.aio import DefaultAzureCredential

@asynccontextmanager
async def connect_voicelive_sdk(*, settings):
    credential = DefaultAzureCredential()
    try:
        # SDK reshapes https://.services.ai.azure.com/ →
        # wss://.services.ai.azure.com/voice-live/realtime
        # ?api-version=2026-04-10&model=
        async with connect(
            credential=credential,
            endpoint=settings.azure_voice_live_endpoint,   # https://, NOT wss://
            api_version="2026-04-10",                       # default in 1.2.0
            model=settings.azure_deployment_name,           # e.g. "gpt-realtime"
            # credential_scopes default = ["https://ai.azure.com/.default"]
        ) as conn:
            await conn.session.update(session=RequestSession(
                modalities=[Modality.TEXT, Modality.AUDIO],
                instructions="You are a friendly assistant.",
                voice=AzureStandardVoice(name="en-US-Ava:DragonHDLatestNeural"),
                input_audio_format=InputAudioFormat.PCM16,
                output_audio_format=OutputAudioFormat.PCM16,
                turn_detection=AzureSemanticVad(
                    create_response=True,        # §12.1 proactive
                    auto_truncate=True,          # §12.4 token governance
                ),
            ))
            yield conn
    finally:
        await credential.close()

Why move to Rung 4

| Concern | Rungs 1–3 (openai shim) | Rung 4 (azure-ai-voicelive) | |---------|---------------------------|-------------------------------| | Typed session config | dict literals | RequestSession + typed models | | Endpoint shape | wss://…/voice-live + base override | https://…services.ai.azure.com (SDK derives) | | Auth scope default | manual https://ai.azure.com/.default | SDK default https://ai.azure.com/.default | | MCP tools mid-turn | manual JSON | MCPServer + MCPTool typed | | Avatar / custom voice | manual JSON | AvatarConfig, AzureCustomVoice | | Interim response | not exposed | LlmInterimResponseConfig | | API version pin | env var | SDK constant (override via kwarg) |

The native SDK is the recommended path for new code as of GA 2026-04-10. Migrate Rungs 1–3 incrementally — the wire protocol is identical, so the audio pipe and event handler can stay as-is.

Endpoint hostname (services.ai vs cognitiveservices)

Voice Live lives on the services.ai.azure.com subdomain of your Foundry resource — the SAME resource that serves chat models on cognitiveservices.azure.com. Map your CI/prod env var like:

endpoint = os.environ["AZURE_AI_ENDPOINT"].replace(
    "cognitiveservices.azure.com", "services.ai.azure.com"
)
# Or set AZURE_VOICELIVE_ENDPOINT directly to the services.ai host.

> Install: pip install "azure-ai-voicelive[aiohttp]~=1.2.0". > The [aiohttp] extra is required for the async connect > path — without it the import raises RuntimeError: aiohttp not > installed.


3 · Session Configuration

Voice Live sessions expose capabilities that plain Realtime doesn't. Send these in conn.session.update(session={...}) after connection.

Realtime session (Rung 1)

{
    "turn_detection": {"type": "server_vad"},
    "input_audio_format": "pcm16",
    "output_audio_format": "pcm16",
    "voice": "alloy",                              # OpenAI voice set only
    "instructions": "You are a friendly assistant.",
    "modalities": ["text", "audio"],
    "input_audio_transcription": {
        "model": "whisper-1",
        "language": "en",
    },
}

Voice Live session (Rung 2)

{
    "turn_detection": {"type": "azure_semantic_vad", "remove_filler_words": False},
    "input_audio_format": "pcm16",
    "output_audio_format": "pcm16",
    "voice": {"name": "en-US-Ava:DragonHDLatestNeural", "type": "azure-standard"},
    "instructions": "You are a friendly assistant.",
    "modalities": ["text", "audio"],
    "input_audio_echo_cancellation": {"type": "server_echo_cancellation"},
    "input_audio_noise_reduction": {"type": "azure_deep_noise_suppression"},
    "input_audio_transcription": {
        "model": "azure-fast-transcription",        # faster than whisper-1
        "language": "en",
    },
}

Voice Live + Agent session (Rung 3)

The agent owns instructions and tools — omit instructions from the session config:

{
    "turn_detection": {"type": "azure_semantic_vad", "remove_filler_words": False},
    "input_audio_format": "pcm16",
    "output_audio_format": "pcm16",
    "voice": {"name": "en-US-Ava:DragonHDLatestNeural", "type": "azure-standard"},
    "modalities": ["text", "audio"],
    "input_audio_echo_cancellation": {"type": "server_echo_cancellation"},
    "input_audio_noise_reduction": {"type": "azure_deep_noise_suppression"},
    "input_audio_transcription": {
        "model": "azure-fast-transcription",
        "language": "en",
    },
}

Non-English semantic VAD

azure_semantic_vad is English-tuned. For other languages, the endpointer can fire mid-utterance on natural hesitations ("uhm…", a pause before an ID), clipping the user. Two fixes:

  • Use the multilingual VAD variant and a multilingual end-of-utterance

detection model for non-English locales.

  • silence_duration_ms is the lever that bridges mid-sentence pauses

threshold and timeout_ms alone don't. Raising it tolerates longer pauses at the cost of a little latency per turn (≈ the extra silence you wait for), so tune it to the locale's natural pausing, not lower.

Feature comparison

| Feature | Realtime | Voice Live | |---------|----------|------------| | VAD | server_vad | azure_semantic_vad (understands pauses vs hesitation) | | Echo cancellation | ❌ | server_echo_cancellation (built-in AEC) | | Noise reduction | ❌ | azure_deep_noise_suppression | | Transcription | whisper-1 | azure-fast-transcription (lower latency) | | Voice set | OpenAI only (10 voices) | Azure Neural HD + OpenAI (per locale) | | Voice format | bare string "alloy" | {"name": "...", "type": "azure-standard"} | | Filler word removal | ❌ | Optional (remove_filler_words: true) |

Cascade vs native realtime — latency facts

A native realtime model (gpt-realtime*) speaks directly. A text model on Voice Live (gpt-4o, gpt-5*, …) runs as a cascade: the model emits text, then a managed Azure TTS overlay speaks it. Two consequences worth designing for:

  • First-audio floor. Audio can't start before the model's first token, so the

cascade has an inherent first-audio floor that no client tuning removes. Native realtime has no overlay and starts sooner. Pick native realtime when first-audio latency is the priority; pick the cascade when you need a specific text model's reasoning or a managed model with no deployment.

  • **reasoning_effort is not monotonic for perceived latency.** Lower effort

can change behavior, not just speed: minimal may reduce orchestration quality or make a model skip a preamble/acknowledgment turn. A middle setting is a reasonable starting point, but the effect is model-, prompt-, and tool-flow- dependent — measure per scenario with the two-track method in §7 rather than assuming "lower effort = faster experience".


4 · Voice Catalog

OpenAI voices (all rungs)

All 10 are locale-independent: alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar.

On Realtime (Rung 1), send as a bare string: "voice": "alloy". On Voice Live (Rung 2–3), wrap with type: "voice": {"name": "alloy", "type": "azure-standard"}.

Azure Neural HD voices (Voice Live only)

Per-locale catalog. The type field is always "azure-standard".

English:

| Voice | Name string | |-------|-------------| | Ava (default) | en-US-Ava:DragonHDLatestNeural | | Jenny | en-US-Jenny:DragonHDLatestNeural | | Davis | en-US-Davis:DragonHDLatestNeural | | Guy | en-US-GuyNeural | | Brian | en-US-BrianNeural |

Italian:

| Voice | Name string | |-------|-------------| | Isabella (default) | it-IT-IsabellaMultilingualNeural | | Giuseppe | it-IT-GiuseppeMultilingualNeural | | Alessio | it-IT-AlessioMultilingualNeural | | Marta | it-IT-MartaNeural | | Diego | it-IT-DiegoNeural | | Elsa | it-IT-ElsaNeural |

> Localization note: The upstream demo ships English (en) and Italian > (it) locale packs. Other locales follow the same pattern — one entry > per dict in i18n.py.

Voice fallback on rung switch

When switching from Voice Live → Realtime at runtime, HD voice names are not valid for the Realtime API. The handler falls back to alloy if the current voice isn't in the OpenAI set — the UI also snaps the picker to a valid value (belt-and-braces).

Avatars + custom voice

> 🆕 //build 2026 — public preview.

Voice Live now ships avatar rendering (lip-sync video over the audio stream) and custom voice (BYO speaker model registered via Azure Speech Studio). Both surfaces live behind the same WSS connection — no new SDK; toggle via session.update extensions. Configure in the Foundry portal under your Voice Live resource. The STT/TTS locale catalog grew to 140+ STT / 600+ TTS at GA, so the existing voice picker in ui/voice_picker.py will auto-populate any new options surfaced by /voices.


5 · Foundry Agent Routing

Rung 3 routes the Voice Live WebSocket to a Foundry Agent. The agent can be a hosted agent (MAF container) or a prompt agent (declarative). The routing is done via extra_query:

extra_query={
    "agent-id":           "",            # from Foundry portal
    "agent-project-name": "",        # Foundry project name
    "agent-access-token": await token_provider(),  # ai.azure.com scope
}

Two token scopes

| Scope | Used for | |-------|----------| | https://cognitiveservices.azure.com/.default | Model access (Realtime + Voice Live) | | https://ai.azure.com/.default | Foundry Agent routing (Rung 3 only) |

from azure.identity.aio import DefaultAzureCredential, get_bearer_token_provider

credential = DefaultAzureCredential()

# Model scope — all

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [aiappsgbb](https://github.com/aiappsgbb)
- **Source:** [aiappsgbb/awesome-gbb](https://github.com/aiappsgbb/awesome-gbb)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.