Install
$ agentstack add skill-docker-skills-docker-agent-deploy ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
Docker Agent: Serving, Sharing, and Evaluating
Overview
This skill owns the integration surface of Docker Agent: making an agent reachable by other software (docker agent serve), distributing it through an OCI registry the way container images are distributed (docker agent share), and proving it still behaves after a change (docker agent eval). It assumes the agent config already exists — see docker-agent-config for authoring it, and docker-agent-run for interactive/local invocation.
When to use this skill
Activate this skill when:
- The user wants an agent reachable over MCP, an OpenAI-compatible chat endpoint, a plain HTTP API, or A2A/ACP.
- The user wants to publish an agent to Docker Hub (or any OCI registry) or pull one someone else published.
- The user wants automated evaluations (regression tests) for an agent, or wants to gate CI on eval results.
Do not use this skill when
Do not use this skill when:
- The task is authoring the agent.yaml itself (models, toolsets, sub_agents) — use
docker-agent-config. - The task is running the agent interactively on a developer's machine, choosing
--safety/--sandbox, or aliases — usedocker-agent-run.
Core guidance
Serving an agent
- Five server modes, each with its own default loopback listen address —
never expose any of them beyond loopback without authentication:
| Mode | Default listen | Auth flag | Has --safety? | | --- | --- | --- | --- | | serve mcp | 127.0.0.1:8081 | --auth-token (only with --http) | Yes (only with --http) | | serve api | 127.0.0.1:8080 | --auth-token | No | | serve chat | 127.0.0.1:8083 | --api-key / --api-key-env | Yes | | serve a2a | 127.0.0.1:8082 | --auth-token | Yes | | serve acp | (stdio only) | n/a | No |
``bash docker agent serve mcp ./agent.yaml --http --listen 127.0.0.1:9090 --auth-token "$TOKEN" ``
serve mcpdefaults to stdio transport (for local clients like Claude
Desktop); pass --http only when you need a network-reachable MCP endpoint, and set --auth-token whenever you do.
- Binding any server flag to a non-loopback address without an auth
token/key is refused; --insecure-no-auth exists to force it and must be treated as a deliberate, documented exception, never a default.
serve mcp(with--http),serve chat, andserve a2aexpose
--safety (strict/balanced/restricted/autonomous); Docker's docs state it defaults to restricted for these modes when unset. serve api and serve acp expose no --safety flag at all. Never raise --safety to autonomous on a network-reachable listener; if a served agent must approve more, prefer balanced and keep auth enabled.
serve apiaccepts a directory instead of a single file: every
.yaml/.yml/.hcl in it is exposed under /api/agents. Use --session-workingdir-root to confine session working directories when the server is reachable by more than one user.
Sharing agents via OCI registries
- Push and pull agent configs the same way you push and pull images — same
registry, same docker login auth: ``bash docker agent share push ./agent.yaml docker.io/username/my-agent:latest docker agent share pull docker.io/username/my-agent:latest ``
instruction_filecontents are inlined into the pushed artifact
automatically, so a published agent stays self-contained — you do not need to bundle the referenced files separately.
- Pin
sub_agentsthat reference the pushed artifact to a digest
(name@sha256:...) once published, to avoid a per-run registry lookup and to guarantee the exact config a consumer gets.
- Use
--forceonshare pullonly when you intend to overwrite a local
copy that already exists; without it, an existing local config is left untouched.
Evaluating agents
- Evals live in an
evals/directory next to the agent config by default;
each eval is one JSON session file capturing a user message, the recorded tool calls, and an evals object with the scoring criteria.
- Create eval sessions from real conversations rather than hand-writing
JSON: run the agent interactively, then use the /eval slash command in the TUI to save the session, and edit in relevance/size/assertions criteria afterward.
- Four scoring dimensions: Tool Calls (F1 against the recorded sequence),
Relevance (LLM-judge, --judge-model, default anthropic/claude-opus-5), Size (S/M/L/XL response-length bucket), and Assertions (deterministic checks; see the complete assertion-type list in references/eval-format.md). Prefer assertions over relevance when a check can be exact: they need no judge model and are deterministic, not approximation-prone.
- Evaluations run inside containers for isolation; a Docker-compatible
runtime is required. Dedicated provider API keys (ANTHROPIC_API_KEY/OPENAI_API_KEY) are forwarded automatically. GITHUB_TOKEN/GH_TOKEN are not forwarded automatically (they're broad host credentials, not model keys) — pass them explicitly with -e GITHUB_TOKEN when an agent's provider needs one (e.g. github-copilot).
- Gate CI on regressions, not on absolute scores, with
--baseline:
``bash docker agent eval ./agent.yaml --baseline results/2026-08-01-run.json --regression-tolerance 0.05 ` A previously-passing eval that now fails always gates regardless of tolerance; cost changes are reported but never gate. A baseline or run with zero evaluations (e.g. an --only` pattern matching nothing) is rejected rather than reported as passing.
- Use
--keep-containersplus your runtime'sexecto inspect a failed
eval's container; the eval's .db session file holds the full conversation for offline debugging.
Verify
- After changing a served agent's config, re-run its evals with the same
explicit --safety value used in the deployment before restarting the listener — this catches an approval-policy regression before it reaches traffic. If a rollout must be rolled back, restore the prior config and safety flag; never restore an unauthenticated listener as a rollback shortcut.
Related skills
- For writing or changing the underlying
agent.yaml, usedocker-agent-config. - For local/interactive runs, safety-mode choice, and sandboxing, use
docker-agent-run.
References
references/eval-format.md— full eval session JSON schema and CLI flag table.references/sources.md— provenance of every rule in this skill.
Assets
assets/eval-session-example.json— a minimal eval session file to copy and adapt.
Checks
checks/verification.md— Verification runbook for serving, sharing, and evaluating an agent.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: docker
- Source: docker/skills
- License: Apache-2.0
- Homepage: https://docs.docker.com/ai/skills/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.