Install
$ agentstack add mcp-mezmo-aura Open-source listing, not yet scanned by AgentStack. Follow the source repository for install instructions.
Security review
⚠ Flagged1 finding(s); flagged for manual review. · v0.1.0 How review works →
- • Prompt-injection patterns
- • Secret / credential exfiltration
- • Dangerous shell & filesystem operations
- • Untrusted network calls
- • Known-malicious package signatures
- high Pipes remote content directly into a shell (remote code execution).
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ● Environment & secrets Used
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
AURA
[](https://mezmo.com/r/slack-aura) [](LICENSE) [](https://www.rust-lang.org) [](https://modelcontextprotocol.io)
AURA is an agentic harness that turns an LLM model into a reliable, autonomous service capable of executing real SRE work. AURA provides the guardrails, API servers, state management, authentication, streaming, error handling, and tool integrations necessary to run AI SRE agents safely in production.
Key capabilities:
- Declarative agent composition via TOML with multi-provider LLM support and multi-agent serving
- Dynamic MCP tool discovery via HTTP streamable, SSE, and STDIO transports
- Automatic schema sanitization for OpenAI function-calling compatibility
- Vector search integration with Qdrant and AWS Bedrock Knowledge Base
- Embeddable Rust core independent from configuration layer
- Multi-agent orchestration with coordinator/worker architecture and DAG-based parallel execution
- Dependency-aware multi-wave execution with plan/execute loops
- A2A protocol support for agent-to-agent interoperability
Table of Contents
- [Quick Start](#quick-start)
- [Project Structure](#project-structure)
- [Development Setup](#development-setup)
- [Usage](#usage)
- [Web API Server](#web-api-server)
- [Client-Side Tools](#client-side-tools)
- [Configuration](#configuration)
- [Multiple Agents](#multiple-agents)
- [Configuration Sections](#configuration-sections)
- [Orchestration](#orchestration)
- [Scratchpad (Context Window Management)](#scratchpad-context-window-management)
- [Ollama](#ollama)
- [Observability](#observability)
- [Development and Testing](#development-and-testing)
- [Testing](#testing)
- [Documentation](#documentation)
- [Architecture](#architecture)
Quick Start
cp .env.example .env # set your LLM provider, model, and API key
docker compose up -d # starts Aura (orchestrator mode) + LibreChat + Phoenix
docker exec -it aura ./aura-cli --api-url http://localhost:8080 # chat with the orchestrator from your terminal
Aura boots in orchestrator mode: a coordinator routes each request — answering simple ones directly and decomposing complex ones across specialized workers. The bundled aura-cli connects to the in-container server automatically and renders the coordinator's plan and worker activity as it streams.
Prefer a browser? Open to chat in LibreChat, or to inspect traces in Phoenix.
[Full quickstart guide](docs/quickstart.md) — provider setup (OpenAI, Anthropic, Ollama, llama-server), adding MCP tools, enabling vector search, serving multiple agents, and troubleshooting.
More Quickstarts
- [Orchestration — Math MCP](examples/quickstart-orchestration-math/README.md) — Multi-agent orchestration with coordinator/worker architecture
- [Kubernetes SRE](examples/quickstart-k8s-sre/README.md) — AI-powered SRE agent on KIND with Kubernetes and Prometheus MCP servers
- [Example Configs](examples/README.md) — Minimal per-provider configs and complete agent compositions
Project Structure
aura/
├── crates/
│ ├── aura/ # Core library (agent builder + orchestration)
│ ├── aura-cli/ # Interactive terminal client (HTTP + standalone modes)
│ ├── aura-config/ # TOML parser and config loader
│ ├── aura-events/ # Shared SSE event types
│ ├── aura-test-utils/ # Shared testing utilities
│ └── aura-web-server/ # OpenAI-compatible HTTP/SSE server
├── compose/ # Docker Compose (integration + orchestration overlays)
├── configs/ # E2E test and orchestration configurations
├── deployment/ # Helm charts and K8s manifests
├── docs/ # Architecture and protocol documentation
├── examples/ # Example and reference configurations
├── scripts/ # CI and utility scripts
└── tests/ # Integration test fixtures and helpers
Development Setup
For building AURA from source without Docker.
- Install Rust if needed:
``bash curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh ``
- Clone and configure:
``bash cd aura cp examples/reference.toml config.toml ``
- Set required environment variables:
``bash export OPENAI_API_KEY="your-api-key" ``
- Build and run:
``bash cargo run --bin aura-web-server ``
Security: keep secrets in environment variables and reference them in TOML using {{ env.VAR_NAME }}.
Usage
Web API Server
Run the web server:
# Default: reads config.toml
cargo run --bin aura-web-server
# Custom config file
CONFIG_PATH=my-config.toml cargo run --bin aura-web-server
# Config directory (serves multiple agents)
CONFIG_PATH=configs/ cargo run --bin aura-web-server
# Host/port override
HOST=0.0.0.0 PORT=3000 cargo run --bin aura-web-server
# Enable Aura custom SSE events
AURA_CUSTOM_EVENTS=true cargo run --bin aura-web-server
# Kitchen sink: all options
CONFIG_PATH=configs/ \
HOST=0.0.0.0 PORT=8080 \
AURA_CUSTOM_EVENTS=true \
AURA_EMIT_REASONING=true \
TOOL_RESULT_MODE=aura \
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 \
cargo run --bin aura-web-server -- --verbose
Core server options:
| Option | Env Variable | Default | Description | | ---------------------------- | -------------------------- | ------------- | ----------------------------------- | | --config | CONFIG_PATH | config.toml | Path to TOML config file or directory | | --host | HOST | 127.0.0.1 | Bind host | | --port | PORT | 8080 | Bind port | | --server-url | AURA_SERVER_URL | host/port | Canonical public origin published in the A2A agent card (see below) | | --streaming-timeout-secs | STREAMING_TIMEOUT_SECS | 900 | Max SSE request duration | | --first-chunk-timeout-secs | FIRST_CHUNK_TIMEOUT_SECS | 30 | Max time to first provider chunk | | --streaming-buffer-size | STREAMING_BUFFER_SIZE | 400 | SSE backpressure buffer | | --aura-custom-events | AURA_CUSTOM_EVENTS | false | Enable aura.* events | | --aura-emit-reasoning | AURA_EMIT_REASONING | false | Enable aura.reasoning | | --tool-result-mode | TOOL_RESULT_MODE | none | Tool result streaming: none, open-web-ui, aura | | --tool-result-max-length | TOOL_RESULT_MAX_LENGTH | 1000 | Max chars before truncation (aura events) | | --shutdown-timeout-secs | SHUTDOWN_TIMEOUT_SECS | 30 | Graceful shutdown window |
Tool result modes:
none: spec-compliant; tool results appear only in model summary.open-web-ui: tool results emitted throughtool_callsfor OpenWebUI compatibility.aura: tool results emitted viaaura.tool_completeevents.
API examples:
# Health
curl http://localhost:8080/health
# List available models (agents)
curl http://localhost:8080/v1/models
# OpenAI-compatible chat completion
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Hello"}]}'
# Select a specific agent by name or alias via the model field
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "my-agent", "messages": [{"role": "user", "content": "Hello"}]}'
# Streaming response
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Hello"}], "stream": true}'
SSE protocol details, event types, custom events, and client handling are documented in [docs/streaming-api-guide.md](docs/streaming-api-guide.md).
A2A Protocol
> Disabled by default. A2A endpoints are only activated when the server is started with --enable-a2a (or AURA_ENABLE_A2A=true). Omitting the flag means no A2A routes are registered and the agent card is not served.
Aura exposes A2A protocol endpoints for agent-to-agent interoperability. This allows other A2A-compatible agents and clients to discover and interact with Aura agents using a standardized protocol.
# Agent card (capability discovery)
curl http://localhost:8080/.well-known/agent-card.json
# Send a message via REST
curl -X POST http://localhost:8080/a2a/v1/message:send \
-H "Content-Type: application/json" \
-H "A2A-Version: 1.0" \
-d '{"message": {"messageId": "msg-001", "role": "ROLE_USER", "parts": [{"text": "Hello"}]}}'
# Send a message via JSON-RPC
curl -X POST http://localhost:8080/a2a/v1/rpc \
-H "Content-Type: application/json" \
-d '{"jsonrpc": "2.0", "method": "SendMessage", "params": {"message": {"messageId": "msg-002", "role": "ROLE_USER", "parts": [{"text": "Hello"}]}}, "id": 1}'
> Set AURA_SERVER_URL when running behind a proxy, load balancer, or in Kubernetes. The agent card must advertise absolute endpoint URLs, and A2A clients use those URLs directly — a relative or wrong-host URL makes message:send fail even though the card itself loads. Aura builds the card's URLs from AURA_SERVER_URL (or --server-url); set it to the externally-reachable origin clients use (e.g. https://aura.example.com). When unset, it falls back to the bind host/port, which is only correct for direct local access.
A2A endpoints, transport modes, the agent card URL, task lifecycle, and testing examples are documented in [docs/a2a-implementation.md](docs/a2a-implementation.md).
Client-Side Tools
> --- > # USE AT YOUR OWN RISK > --- > > **Setting enable_client_tools = true on an agent grants the LLM the ability to call tools that execute on the client's machine. When clients (e.g. aura-cli) advertise tools like Shell, Read, or Update, the LLM can invoke them and the client will execute them with the privileges of the user running the client. This is functionally equivalent to giving the model a shell prompt on every connecting client. > > The risks are real: > - Prompt injection. Anything the model reads — a file, an MCP tool output, a vector-store hit, a URL — can contain instructions that hijack the model into running destructive commands. The server cannot tell a legitimate request from an injected one. > - Hallucination. The model can confidently call the wrong tool with the wrong arguments. There is no undo for a Shell("rm -rf ...") invocation. > - No server-side sandbox. The server only forwards tool calls; execution happens client-side with full host privileges. Whatever sandboxing exists is the client's responsibility. > - Per-agent permission filters reduce blast radius but are not a security boundary.** client_tool_filter controls which tools the model can ask for, not what they do once invoked. > > Only enable on agents where: > - You trust the model, the provider, and every data source the model can read (configs, MCP servers, vector stores, web fetches). > - You trust every client that will connect with --enable-client-tools and the user account it runs under. > - You and your users accept that worst-case loss (deleted files, leaked credentials, modified source) is acceptable or recoverable. > > Disabled by default. Opting an agent in is your decision and your responsibility — and your users'. > > See [aura-cli's matching warning](crates/aura-cli/README.md#client-side-tools) for the client-side perspective.
> Single-agent configurations only. Client-side tools are not supported in orchestrated (multi-agent) configurations — when [orchestration].enabled = true, any tools array on the request is dropped with a warning. The reason: the passthrough mechanism requires terminating the user-facing SSE stream with finish_reason: "tool_calls", which doesn't compose with the coordinator/worker pipeline. If you need local tools, use a single-agent config.
The server honors a tools array on incoming chat completion requests. Whether those tools are actually attached to the LLM is a per-agent opt-in in TOML — there is no server-wide flag. Tools that get attached are registered as passthrough tools: the LLM sees them alongside any server-side MCP tools and can call them, but instead of executing server-side, the stream terminates with finish_reason: "tool_calls" so the client can run the tool locally and submit the result back as a role: "tool" follow-up.
[agent]
name = "Assistant"
system_prompt = "..."
enable_client_tools = true
client_tool_filter = ["Read", "ListFiles", "Find*"] # optional; omitted/empty = all
client_tool_filter is a list of glob patterns matched against the request's tools[].function.name. An empty or omitted filter means all client tools are available. A request that supplies tools never reaches an agent that did not opt in.
# 1) Initial request advertising a client-side tool
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"stream": true,
"messages": [{"role": "user", "content": "What time is it?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_current_time",
"description": "Get the current time",
"parameters": {"type": "object", "properties": {}}
}
}]
}'
# 2) Stream ends with finish_reason: "tool_calls". The client executes the tool
# locally and submits the result back in a follow-up request:
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"stream": true,
"tools": [ ... same tools array ... ],
"messages": [
{"role": "user", "content": "What time is it?"},
{"role": "assistant", "content": null, "tool_calls": [
{"id": "call_abc", "type": "function",
"function": {"name": "get_current_time", "arguments": "{}"}}
]},
{"role": "tool", "tool_call_id": "call_abc", "content": "2026-04-30T14:30:00Z"}
]
}'
When the loaded agent doesn't opt in (the default), any tools field on the request is silently dropped; the server runs MCP tools as usual but never asks the client to execute anything. Per-agent opt-in is the design — accepting client-supplied tool definitions means trusting the client to execute them, so it should be a deliberate config decision. See [aura-cli](crates/aura-cli/README.md#client-side-tools) for the matching client-side flag (--enable-client-tools) and how the two halves coordinate.
Configuration
Recent breaking changes
- 21 April 2026:
[llm]moved under[agent.llm]; workers may override via[orchestration.worker..llm]. See [migration guide](docs/breaking-changes/20260421-llm-under-agent.md). - 10 April 2026: Several fields moved from
[agent]to[llm]; Ollama params consolidated under[llm.additional_params]. See [migration guide](docs/breaking-changes/20260410-agent-llm-toml-configuration.md).
CONFIG_PATH can point to a single TOML file or a directory of .toml files. When pointed at a directory, AURA loads every .toml file and serves each as a selectable agent. Clients choose an agent via the model field in chat completion requests — the same field that tools like LibreChat, OpenWebUI, and CLI clients use to present a model picker.
Multiple Agents
To serve multiple agents, create a directory with one TOML file per agent:
configs/
├── research-assistant.toml
├── devops-agent
…
## Source & license
This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.
- **Author:** [mezmo](https://github.com/mezmo)
- **Source:** [mezmo/aura](https://github.com/mezmo/aura)
- **License:** Apache-2.0
- **Homepage:** https://www.mezmo.com/aura
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.