AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
MCP verified MIT Self-run

Capgate

mcp-jersonboydmilan-capgate · by jersonboydmilan

Capgate — runtime authority boundary for autonomous AI agents: the agent proposes, the harness authorizes, the executor acts. Deny-by-default policy, non-transitive delegation, tamper-evident audit, MCP gateway, containerised isolation.

— No reviews yet
0 installs
25 views
0.0% view→install

Install

$ agentstack add mcp-jersonboydmilan-capgate

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ● Network access Used
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ✓ Environment & secrets No
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/mcp-jersonboydmilan-capgate)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● 7d ago

Declared compatibility

Claude CodeClaude DesktopCursorWindsurf

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Capgate? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Capgate

Authority management for autonomous software agents.

The agent proposes. The harness authorizes. The executor acts.

Capgate is a runtime enforcement boundary that sits underneath your existing agent stack. Every consequential action an agent takes — a tool call, a message to another agent, a request that another agent act — is a structured proposal that passes through one deterministic, deny-by-default policy engine before anything executes. Every decision is written to a tamper-evident audit trail.

It is not an agent framework, a prompt library, or a model wrapper. It does not ask the model to follow rules; it makes unauthorized actions structurally unable to reach a tool.

Contents

  • [The boundary, under attack](#the-boundary-under-attack) — a compromised agent, contained
  • [Inspect](#inspect-simulation-audit-escalations-policy) — the local UI (screenshots)
  • [The contract](#the-contract) · [The API](#the-api) · [Delegation](#delegation-does-not-transfer-authority)
  • [What is actually enforced](#what-is-actually-enforced) · [MCP](#mcp-agents-call-in-through-the-boundary)
  • [Audit trail](#audit-trail) · [Tests](#tests) · [Status](#status)
  • Design: [DESIGN.md](DESIGN.md) (the five invariants) · [Workload identity](docs/workload-identity.md) · [Isolation runtime](docs/isolation.md) · [State & multi-host](docs/state.md) · [docs/threat-model.md](docs/threat-model.md) · [Roadmap](docs/roadmap.md) · [SECURITY.md](SECURITY.md)

The boundary, under attack

Run a compromised agent — one that ignores the SDK and tries every route around the harness — inside the reference isolated runtime:

cd deploy
python fetch_artifacts.py               # once; images build offline
docker compose --profile agent build     # rebuild every image, including netguard and agent
docker compose up -d harness
docker compose run --rm agent           # the compromised agent, inside its sandbox
docker compose logs harness             # every decision; one successful execution

or, with all checks evaluated automatically:

python deploy/demo.py                   # ~1 minute
CAPGATE — COMPROMISED AGENT IN ISOLATED RUNTIME

  PASS  HTTP  direct call to tool endpoint (no auth)                        0
  PASS  HTTP  out-of-contract action via harness                            [403, 'TOOL_NOT_ALLOWED']
  PASS  HTTP  impersonate db-admin                                          [403, 'IDENTITY_MISMATCH']
  PASS  HTTP  claim admin contract                                          [403, 'CONTRACT_MISMATCH']
  PASS  HTTP  approve own escalation                                        404
  PASS  HTTP  delegate to privileged agent                                  [403, 'TOOL_NOT_ALLOWED']
  PASS  CRED  sub swapped to db admin                                       401
  PASS  CRED  revoked token                                                 401
  PASS  NET   tcp tools:9100 by IP                                          error:101
  PASS  NET   misattached tool on the agent's own network                   refused
  PASS  NET   egress allowlist installed (default DROP, only harness:8080)  iptables OUTPUT
  PASS  NET   internet tcp 1.1.1.1:443                                      error:101
  PASS  NET   resolve example.com                                           dns_failed
  PASS  NET   raw socket                                                    permission_denied
  PASS  PROC  effective capabilities                                        0
  PASS  SEC   filesystem + env + /proc scan for harness secrets             none found
  PASS  PATH  in-contract action via harness                                [200, 'succeeded']
  PASS  PATH  side effects at real tool                                     ['/web.search']
  PASS  PATH  audit hash chain intact                                       True
  …
61/61 checks passed. Side effects at the real tool: 1 (the authorized web.search).

What this runtime enforces, precisely: the agent process cannot open a socket to anything except the harness API, and it cannot read the tool credentials, permit signing key or token keyring. Its own credential is a short-lived signed token; tampered, forged, expired, revoked and over-long tokens are rejected. The direct tool calls now fail at the network (0 = no connection), not merely with 401. The only path that produces a side effect is authorize → signed permit → executor. Details, the two network layers, and what is not claimed: [deploy/README.md](deploy/README.md).

Thirty seconds: simulate a contract

pip install -e .
capgate simulate examples/simulation/task.yaml
CAPGATE — SIMULATION

Contract: research-v1
Agent: researcher

    ACTION                          DECISION   REASON
─────────────────────────────────────────────────────
01  web.search                      ALLOW      CAPABILITY_GRANTED
02  web.fetch                       ALLOW      CAPABILITY_GRANTED
03  database.read                   DENY       EXPLICITLY_DENIED
04  agent.delegate                  ESCALATE   REQUIRES_APPROVAL
05  web.search                      ALLOW      CAPABILITY_GRANTED

Summary
────────────────
Allowed:      3
Denied:       1
Escalated:    1

No external actions were executed.

Simulation needs no infrastructure and uses the exact engine that enforcement uses. Write a contract, run your agent's proposals through simulate, tune the contract, then switch to capgate enforce.

Inspect: simulation, audit, escalations, policy

python examples/inspect/demo.py      # live local stack; opens the browser
capgate inspect --task task.yaml --audit audit.jsonl --harness-url http://127.0.0.1:8080 --approver-token-file alice.token

A thin local UI: edit a contract and watch the allow/deny/escalate table change; browse and verify the hash-chained audit trail; approve or reject pending escalations; ask what a hypothetical action would hit and diff two contracts. See [docs/inspect.md](docs/inspect.md).

Each view, full size

Simulate — the contract and the proposed actions side by side; every edit re-runs the decision table.

Audit — the hash-chained trail, integrity checked on load, filterable and correlatable by decision.

Escalations — pending human approvals, with the proposed action, arguments, contract and reason.

Policy — the grants in effect per agent, a "what would this hit?" evaluator, and a contract diff.

The contract

contract_id: research-v1
goal: Research recent work on agent authorization
max_steps: 20
approvers: [alice]
agents:
  researcher:
    capabilities:
      web.search: allow
      web.fetch:
        effect: allow
        constraints:
          allowed_arguments: [url]
          block_private_hosts: true
      database.read: deny
      agent.delegate:
        effect: escalate
        constraints:
          allowed_targets: [writer]
          allowed_actions: [docs.write]

Anything not named is denied. Unknown fields are errors, so a typo cannot quietly widen or drop a rule. See [docs/contracts.md](docs/contracts.md) and [docs/capabilities.md](docs/capabilities.md).

The API

from capgate import Harness, load_contract

harness = Harness(load_contract("contract.yaml"), tools={"web.search": search})

result = harness.authorize(agent="researcher", action="web.search", arguments={"query": "..."})
if result.allowed:
    output = harness.execute(result).output
else:
    print(result.decision.decision, result.reason_code)   # deny TOOL_NOT_ALLOWED

That is most of the surface: authorize, execute, plus send_message, delegate and approve/reject for multi-agent and human-in-the-loop flows. No orchestration system, DSL, model provider, vector database or observability platform is required.

Delegation does not transfer authority

TaskContract A: web.search ✓  web.fetch ✓  database.write ✗  create_agent ✗
TaskContract B: web.search ✓

Agent A → "Ask Agent B to perform database.write"
Harness → evaluates Agent B's contract → DENIED
cd examples/delegation-boundary && python agent_a.py

The authorization question is always who is acting, under which contract, with what capability, on what resource — never who asked. The policy engine has no input through which a delegator could influence the answer. See [docs/delegation.md](docs/delegation.md).

What is actually enforced

The README claims exactly this: all consequential execution is mediated by the harness, in the following precise sense.

| Layer | Enforced by | Tested in | |---|---|---| | Out-of-contract proposals are denied | Pure deny-by-default policy engine | tests/adversarial/*, tests/unit/test_policy.py | | A denied or escalated action never reaches a tool | Executor runs only with a harness-signed, single-use grant bound to the exact agent, action and argument hash | tests/adversarial/unauthorized_tool (Test C) | | An agent process cannot reach the real tool directly | Isolated runtime: agent's only network route is harness:8080 (internal networks + iptables allowlist). Any deployment: tool credentials exist only in the executor and the tool endpoint requires them | tests/adversarial/isolation (pytest -m docker), tests/bypass/test_process_boundary.py | | An agent cannot read harness secrets | Isolated runtime: secrets mounted only into harness/tools, separate PID namespace, no capabilities, read-only filesystem | tests/adversarial/isolation | | An agent cannot impersonate another agent or pick its contract | Identity comes from a verified short-lived signed token (expiry, TTL cap, rotation, revocation), contract from the server-side binding | tests/bypass, tests/unit/test_identity.py, tests/adversarial/scope_expansion | | Restarts don't reset authority | Budgets, used grants, approvals, messages and revocations persist in SQLite | tests/integration/test_persistence.py | | Every decision is auditable | Audit write happens before a decision is returned; failure blocks execution; records are hash-chained | tests/adversarial/audit (Test E) |

The network and secret guarantees hold for the reference runtime in [deploy/](deploy/README.md), or a deployment that reproduces its properties. If an agent instead runs as the same user on the same host as the harness, it may be able to read the harness's memory, files or environment; credential custody still stops direct tool calls, but not that. Kernel and container-runtime escapes are out of scope. The SDK is a convenience, never the boundary. Full detail: [docs/threat-model.md](docs/threat-model.md).

Ways in, one path through

Python SDK (in-process Harness) ─┐
HTTP API (uvicorn, behind nginx) ┤
capgate_client / curl  ──────────┼──► Interceptor ─► Policy ─► Executor ─► Tool / Agent
MCP client → /mcp  ──────────────┘
capgate token keygen --keyring keys.json                                   # rotate later by running it again
capgate serve examples/basic/server.yaml                                   # keyring hot-reloads; state persists
capgate token issue --keyring keys.json --sub researcher --role agent --ttl 15m
capgate token revoke --keyring keys.json --state state.db --token "$TOKEN"
from capgate_client import HarnessClient
client = HarnessClient("http://127.0.0.1:8700", token=supervisor.current_token)  # str or callable
client.act("web.search", {"query": "..."})

Credentials are short-lived HMAC-signed tokens (ah1...) with a verifier-enforced maximum lifetime, key rotation without restart, and revocation. Tokens are issued by the operator or agent supervisor; an agent cannot renew its own.

In production the API runs under uvicorn behind an nginx edge that bounds request sizes, timeouts and paths (see [deploy/](deploy/README.md)); a stdlib transport is kept for development. Both go through one core, fuzzed with Hypothesis.

MCP: agents call in through the boundary

Any MCP client reaches the harness at POST /mcp; each tools/call is authorized under the caller's contract before it runs.

python examples/mcp/demo.py    # drives /mcp with the official MCP SDK client
Tools offered to this agent:  web-search  web-fetch  agent-delegate  capgate-approval_status
  web-search      OK     CAPABILITY_GRANTED
  database-read   ERROR  EXPLICITLY_DENIED
  agent-delegate  ERROR  (requires human approval → approval_id)

Identity comes from the bearer token, never the MCP payload. Denied and unknown tools are audited tool-errors; escalations return an approval id. Upstream MCP servers can back executor tools with their credentials held harness-side, and capgate mcp-bridge adapts stdio-only clients. See [docs/mcp.md](docs/mcp.md).

A real, model-driven agent loop — Claude choosing tool calls, Capgate authorizing each one — is in [examples/agent_loop](examples/agent_loop/README.md) (python examples/agent_loop/demo.py; runs offline, or drives a real Claude model with ANTHROPIC_API_KEY).

Audit trail

capgate enforce examples/basic/task.yaml --audit audit.jsonl
capgate audit audit.jsonl --decision deny
capgate audit audit.jsonl --verify

Each record states what was proposed, by which agent, under which contract (id and content hash), which capability and policy rule applied, the decision and reason code, and the outcome. Nothing from model reasoning is captured.

Tests

pip install -e ".[dev]" && pytest     # unit, integration, adversarial, fuzz, MCP (~1 min)
pytest -m docker                      # isolated runtime acceptance tests (Docker)
HYPOTHESIS_PROFILE=deep pytest tests/fuzz   # thousands of examples per property

The suite runs on both HTTP transports; [CI](.github/workflows/ci.yml) runs Python 3.10–3.13 on both, plus the Docker isolation job and a nightly deep-fuzz run. Trying to break the boundary is the most useful contribution — see [SECURITY.md](SECURITY.md).

The adversarial suite is organised by attack category — unauthorized_tool, argument_violation, scope_expansion, delegation, capability_expiration, contract_tampering, message_injection, bypass_attempt, budget_exhaustion, escalation, audit — plus tests/adversarial/test_required_scenarios.py, which covers the six acceptance scenarios one test each; tests/bypass/, which attacks a live deployment from a separate OS process; tests/adversarial/isolation/, which attacks the containerised reference deployment from inside the agent's sandbox; and tests/fuzz/, property-based fuzzing of the API, policy engine and both HTTP transports.

Layout

src/harness/      contract, capability, policy, decision, request, interceptor,
                  executor, audit, core (Harness), simulation, cli, tools, identity,
                  state, ratelimit, api (transport-independent core), asgi + server
                  (uvicorn/stdlib transports), inspect/ (local UI), mcp/ (gateway,
                  upstream, stdio bridge)
sdk/python/       capgate_client — thin HTTP client
examples/         basic, simulation, delegation-boundary, adversarial-agent, inspect, mcp, agent_loop
deploy/           reference isolated deployment: compose, harness/, agent/, network/, proxy/ (nginx), gvisor/, k8s/, demo
policies/         reusable contract templates
docs/             architecture, contracts, capabilities, delegation, simulation, mcp, inspect, isolation, threat model
.github/          CI: 3.10–3.13 × both transports, k8s manifests, docker + gVisor isolation, nightly deep fuzz
DESIGN.md         the five invariants  ·  SECURITY.md  ·  CONTRIBUTING.md

Status

The five pieces — contract, proposal, deterministic decision, enforced execution, audit — work end to end and are tested together (tests/integration/test_end_to_end.py). In the reference isolated deployment the boundary holds against every tested bypass route, enforced by the network and process boun

…

Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.