Install
$ agentstack add skill-cartesia-ai-skills-line-voice-agent Open-source listing — not yet scanned by AgentStack. Follow the source repository for install instructions.
Security review
⚠ Flagged1 finding(s); flagged for manual review. · v0.1.0 How review works →
- • Prompt-injection patterns
- • Secret / credential exfiltration
- • Dangerous shell & filesystem operations
- • Untrusted network calls
- • Known-malicious package signatures
- high Pipes remote content directly into a shell (remote code execution).
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ● Environment & secrets Used
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Line SDK Voice Agent Guide
Build production voice agents with the Cartesia Line SDK. This guide covers agent creation, tool patterns, multi-agent workflows, and LLM provider configuration.
How Line Works
Line is Cartesia's voice agent deployment platform. You write Python agent code using the Line SDK, deploy it to Cartesia's managed cloud via the cartesia CLI, and Cartesia hosts it with auto-scaling. Cartesia handles STT (Ink), TTS (Sonic), telephony, and audio orchestration. Only one deployment per agent is active at a time; once deployed, your agent receives calls automatically.
┌─────────────────────────────────────────────────────────────────┐
│ Cartesia Line Platform │
│ ┌──────────┐ ┌──────────────┐ ┌──────────┐ │
│ │ Ink │───▶│ Your Agent │───▶│ Sonic │ │
│ │ (STT) │ │ (Line SDK) │ │ (TTS) │ │
│ └──────────┘ └──────────────┘ └──────────┘ │
│ ▲ │ │
│ │ Audio Orchestration │ │
│ └────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
▲ │
│ WebSocket ▼
┌───────┴────────────────────────────────────┴───────┐
│ Client (Phone / Web / Mobile) │
└─────────────────────────────────────────────────────┘
Your code handles:
- LLM reasoning and conversation flow
- Tool execution (API calls, database lookups)
- Multi-agent coordination and handoffs
Cartesia handles:
- Speech-to-text (Ink)
- Text-to-speech (Sonic)
- Real-time audio streaming
- Turn-taking and interruption detection
- Deployment and auto-scaling
Audio Input Options:
- Cartesia Telephony - Managed phone numbers
- [Calls API](references/calls-api.md) - Web apps, mobile apps, custom telephony
Prerequisites
- Python 3.10+ and uv (recommended package manager)
- Cartesia API key — get one at play.cartesia.ai/keys (used by the CLI and for deployment)
- LLM API key — for whichever LLM provider your agent calls (e.g.
ANTHROPIC_API_KEY,OPENAI_API_KEY,GEMINI_API_KEY) - Cartesia CLI — install with:
``bash curl -fsSL https://cartesia.sh | sh ``
Cartesia CLI Reference
# Authentication
cartesia auth login # Login with Cartesia API key
cartesia auth status # Check auth status
# Project Setup
cartesia create [project-name] # Create project from template
cartesia init # Link existing directory to an agent
# Local Development
cartesia chat # Chat with local agent (text mode)
# Deployment
cartesia deploy # Deploy to Cartesia cloud
cartesia status # Check deployment status
# Environment Variables (encrypted, stored on Cartesia)
cartesia env set KEY=VALUE # Set a single env var
cartesia env set --from .env # Import all vars from .env file
cartesia env rm # Remove an env var
# Agents & Calls
cartesia agents ls # List all agents
cartesia deployments ls # List deployments
cartesia call [agent-id] # Make outbound call
Full command reference: docs.cartesia.ai/line/cli.
Quick Start
1. Create Project
cartesia auth login
cartesia create my-agent
cd my-agent
2. Write Agent Code
main.py:
import os
from line.llm_agent import LlmAgent, LlmConfig, end_call
from line.voice_agent_app import AgentEnv, CallRequest, VoiceAgentApp
async def get_agent(env: AgentEnv, call_request: CallRequest):
return LlmAgent(
model="anthropic/claude-haiku-4-5-20251001",
api_key=os.getenv("ANTHROPIC_API_KEY"),
tools=[end_call],
config=LlmConfig(
system_prompt="You are a helpful voice assistant.",
introduction="Hello! How can I help you today?",
),
)
app = VoiceAgentApp(get_agent=get_agent)
if __name__ == "__main__":
app.run()
3. Test Locally
ANTHROPIC_API_KEY=your-key python main.py
cartesia chat 8000 # Text chat with your running agent
4. Deploy
cartesia env set ANTHROPIC_API_KEY=your-key # Encrypted, stored on Cartesia
cartesia deploy
cartesia status # Verify deployment is active
5. Make a Call
cartesia call +1234567890 # Outbound call via CLI
Or trigger calls from the Cartesia dashboard.
Project Structure
Every Line agent project MUST have:
my_agent/
├── main.py # VoiceAgentApp entry point (REQUIRED)
├── cartesia.toml # Deployment config, created by cartesia init or cartesia create (REQUIRED)
└── pyproject.toml # Dependencies: cartesia-line
cartesia.toml declares deployment metadata, the local server address, and the env vars your agent requires:
[cartesia]
name = "My Agent"
description = "What this agent does"
version = "0.1.0"
[cartesia.server]
port = 8000
host = "0.0.0.0"
[cartesia.environment]
required_vars = ["ANTHROPIC_API_KEY"]
Core Concepts
LlmAgent
The main agent class that wraps LLM providers via LiteLLM:
from line.llm_agent import LlmAgent, LlmConfig
agent = LlmAgent(
model="gemini/gemini-2.5-flash-preview-09-2025", # LiteLLM model string
api_key=os.getenv("GEMINI_API_KEY"), # Provider API key
tools=[end_call, my_custom_tool], # List of tools
config=LlmConfig(...), # Agent configuration
max_tool_iterations=10, # Max tool call loops (default: 10)
backend=None, # Optional provider backend override
)
LlmConfig
Configuration for agent behavior and LLM sampling:
from line.llm_agent import LlmConfig
config = LlmConfig(
# Agent behavior
system_prompt="You are a helpful assistant.",
introduction="Hello! How can I help?", # Set to "" to wait for user first
# Sampling parameters (optional)
temperature=0.7,
max_tokens=1024,
top_p=0.9,
stop=["\n\n"],
seed=42,
presence_penalty=0.0,
frequency_penalty=0.0,
# Reasoning models only: "none" | "minimal" | "low" | "medium" | "high"
reasoning_effort="low",
# Resilience (optional)
num_retries=2, # Default: 2
timeout=30.0,
fallbacks=["gpt-5-nano"], # Fallback models
# Advanced (optional)
strict_tool_schemas=True, # Default: True
extra={}, # Provider-specific pass-through kwargs to LiteLLM
)
> reasoning_effort is validated against the model: passing it to a model that > doesn't support reasoning raises ValueError. Use "none" (or omit it) for > non-reasoning models.
Dynamic Configuration from CallRequest
Use LlmConfig.from_call_request() to pull configuration from the incoming call:
async def get_agent(env: AgentEnv, call_request: CallRequest):
return LlmAgent(
model="anthropic/claude-sonnet-4-5",
api_key=os.getenv("ANTHROPIC_API_KEY"),
tools=[end_call],
config=LlmConfig.from_call_request(
call_request,
fallback_system_prompt="Default system prompt if not in request.",
fallback_introduction="Default introduction if not in request.",
temperature=0.7, # Additional LlmConfig options
),
)
Priority order: CallRequest value > fallback argument > SDK default
VoiceAgentApp
The application harness that manages HTTP endpoints and WebSocket connections:
from line.voice_agent_app import VoiceAgentApp, AgentEnv, CallRequest
async def get_agent(env: AgentEnv, call_request: CallRequest):
# env.loop - asyncio event loop
# call_request.call_id - unique call identifier
# call_request.agent.system_prompt - from request
# call_request.agent.introduction - from request
# call_request.metadata - custom metadata dict
return LlmAgent(...)
app = VoiceAgentApp(get_agent=get_agent)
app.run(host="0.0.0.0", port=8000)
Built-in Tools
Import from line.llm_agent:
from line.llm_agent import (
end_call, send_dtmf, transfer_call, web_search,
knowledge_base, mcp_tool, http_server_tool,
)
end_call
End the current call. Tell the LLM to say goodbye before calling this.
tools=[end_call]
# System prompt: "Say goodbye before ending the call with end_call."
send_dtmf
Send DTMF tones (touch-tone buttons). Useful for IVR navigation.
tools=[send_dtmf]
# Buttons: "0"-"9", "*", "#" (strings, not integers!)
transfer_call
Transfer to another phone number (E.164 format required).
tools=[transfer_call]
# Example: +14155551234
web_search
Search the web for real-time information. Uses native LLM web search when available, falls back to DuckDuckGo.
# Default settings
tools=[web_search]
# Custom settings
tools=[web_search(search_context_size="high")] # "low", "medium", "high"
knowledge_base
Look up information from the agent's knowledge base via a natural-language query. Filters, top_k, and timeout_s are fixed at construction time — the LLM only chooses the query string.
# Default behavior — no filters
tools=[knowledge_base]
# Pre-filter every retrieval, override top_k, or run as a background lookup
tools=[knowledge_base(filters={"category": "billing"}, top_k=10)]
tools=[knowledge_base(description="Look up insurance policy terms.")]
tools=[knowledge_base(is_background=True)]
Tell the user you're looking something up before calling it — retrieval can take a moment. Raises KnowledgeBaseError (import from line) on failure.
mcp_tool
Expose a Model Context Protocol server to the LLM. Requires Python 3.10+ and the mcp package (already a Line dependency).
# Remote HTTP/SSE server
tools=[mcp_tool(name="dmcp", server_url="https://dmcp-server.deno.dev/sse")]
# Local stdio server
tools=[mcp_tool(name="memory", command="npx -y @modelcontextprotocol/server-memory")]
The LLM calls the tool with no arguments to list available tools, or with tool_name and tool_args to invoke one.
httpservertool
Create an HTTP/webhook tool from JSON schemas — no custom function needed. The LLM fills in the schema fields and the SDK makes the request. Properties with constant_value are hidden from the LLM and injected into every request; ${ENV_VAR} placeholders in auth are resolved from os.environ at build time.
create_ticket = http_server_tool(
name="create_ticket",
description="Creates a support ticket for the caller.",
url="https://api.example.com/v1/{tenant_id}/tickets", # {param} = path variable
method="POST",
request_body_schema={
"type": "object",
"required": ["subject", "priority"],
"properties": {
"subject": {"type": "string", "description": "Short summary."},
"priority": {"type": "string", "enum": ["low", "medium", "high"]},
"source": {"type": "string", "constant_value": "voice_agent"}, # hidden
},
},
query_params_schema=None, # same shape, scalar types only, for GET query params
auth={"Authorization": "Bearer ${SUPPORT_API_KEY}"},
content_type="application/json", # or "application/x-www-form-urlencoded"
timeout=5.0,
is_background=True, # default True
)
tools=[create_ticket, end_call]
The LLM always receives a structured JSON result, e.g. {"ok": true, "status": 201, "body": "..."} or {"ok": false, "status": 500, "error": "..."}.
> Note: some Line docs/READMEs refer to this as webhook_tool; the exported > function name is http_server_tool.
Custom Tool Types
Three tool paradigms for different use cases:
| Type | Decorator | Use Case | Result Handling | |------|-----------|----------|-----------------| | Loopback | @loopback_tool | API calls, database lookups | Result sent back to LLM | | Passthrough | @passthrough_tool | End call, transfer, DTMF | Bypasses LLM, goes to user | | Handoff | @handoff_tool | Multi-agent workflows | Transfers control to another agent |
Tool Type Decision Tree
Does the result need LLM processing?
├─ YES → @loopback_tool
│ └─ Is it long-running (>1s)? → @loopback_tool(is_background=True)
│ └─ Yield interim status, then final result
├─ NO, deterministic action → @passthrough_tool
│ └─ Yields OutputEvent objects directly (AgentSendText, AgentEndCall, etc.)
└─ Transfer to another agent → @handoff_tool or agent_as_handoff()
Loopback Tools
Results are sent back to the LLM to inform the next response:
from typing import Annotated
from line.llm_agent import loopback_tool, ToolEnv
@loopback_tool
async def get_order_status(
ctx: ToolEnv,
order_id: Annotated[str, "The order ID to look up"],
) -> str:
"""Look up the current status of an order."""
order = await db.get_order(order_id)
return f"Order {order_id} status: {order.status}, ETA: {order.eta}"
Parameter syntax:
- First parameter MUST be
ctx: ToolEnv - Use
Annotated[type, "description"]for LLM-visible parameters - Tool description comes from the docstring
- Optional parameters need default values (not just
Optional[T])
@loopback_tool
async def search_products(
ctx: ToolEnv,
query: Annotated[str, "Search query"],
category: Annotated[str, "Product category"] = "all", # Optional with default
limit: Annotated[int, "Max results"] = 10,
) -> str:
"""Search the product catalog."""
...
Passthrough Tools
Results bypass the LLM and go directly to the user/system:
from line.events import AgentSendText, AgentTransferCall
from line.llm_agent import passthrough_tool, ToolEnv
@passthrough_tool
async def transfer_to_support(
ctx: ToolEnv,
reason: Annotated[str, "Reason for transfer"],
):
"""Transfer the call to the support team."""
yield AgentSendText(text="Let me transfer you to our support team now.")
yield AgentTransferCall(target_phone_number="+18005551234")
Output event types (from line.events):
AgentSendText(text="...")- Speak text to userAgentEndCall()- End the callAgentTransferCall(target_phone_number="+1...")- Transfer callAgentSendDtmf(button="5")- Send DTMF tone
Handoff Tools
Transfer control to another agent. See [Multi-Agent Workflows](references/multi-agent-workflows.md).
Context Management
LlmAgent exposes a history object for injecting and transforming the conversation history the LLM sees.
agent = LlmAgent(model="gemini/gemini-2.5-flash-preview-09-2025", api_key=...)
# Inject a custom entry (defaults to role="user"; pass role="system" for a system note)
agent.history.add_entry("The customer's name is Alice and she has a premium account.")
# Anchor an insertion relative to an existing event
agent.history.add_entry("Reminder: stay concise.", role="system", after=some_event)
# Replace a segment of history with new events (filtering, summarization, etc.)
agent.history.update(new_events, start=first_event, end=last_event)
Entries are inserted lazily and survive across turns. Inside a tool you can call agent.history.add_entry(...) to persist rich context fetched from an external API.
> Note: some Line READMEs show agent.add_history_entry(...) / > agent.set_history_processor(...). The implemented API is agent.history.add_entry(...) > and agent.history.update(...).
Model Selection Strategy
**Use FAST m
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: cartesia-ai
- Source: cartesia-ai/skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.