AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Agent Proof Approval Gate

skill-carloscape-octorato-agent-proof-approval-gate · by CarlosCaPe

Build a fail-closed PreToolUse gate for merge/deploy/destructive actions that the AI agent provably cannot self-bypass. Use when you need human-only override for a consequential action the agent orchestrates.

No reviews yet
0 installs
30 views
0.0% view→install

Install

$ agentstack add skill-carloscape-octorato-agent-proof-approval-gate

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets Used
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-carloscape-octorato-agent-proof-approval-gate)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Agent Proof Approval Gate? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Agent-Proof Approval Gate

Problem

A PreToolUse hook that blocks a destructive action (merge, deploy, delete) needs a human-only override. Naively you might check for an env var or a flag — but the agent can set those itself with an inline prefix (APPROVE=1 gh pr merge 96) or by writing a file. The gate must be unforgeable by the entity it constrains.

Key Insight — Inline Env Never Reaches the Hook

A PreToolUse hook runs in the harness process, not in the shell that executes the agent's command. When the agent writes VAR=1 cmd, that assignment is scoped to the child shell that runs cmd; the hook fires before cmd even starts, in a separate environment. Therefore:

> An env var set in the agent's command prefix is invisible to the hook.

Only a human who runs export OCTO_MERGE_APPROVE=96 in the real terminal session can set the hook's env. The agent cannot reach it.

Design

Primary channel — scoped env var (agent-proof)

# Human grants approval for a specific PR
export OCTO_MERGE_APPROVE=96      # must match the exact PR number being merged

The hook validates:

  1. OCTO_MERGE_APPROVE is set.
  2. Its value equals the PR number extracted from the command being intercepted (not startswith, not in — exact equality).
  3. Optionally, a TTL: compare against the file-mtime of a stamp written when the var was set.
import os, re, sys

def check_approval(pr_number: str) -> bool:
    approved = os.environ.get("OCTO_MERGE_APPROVE", "").strip()
    return approved == pr_number          # "96" != "95", "96x", " 96"

# In the hook body:
if not check_approval(detected_pr):
    print(f"[gate] BLOCKED — set OCTO_MERGE_APPROVE={detected_pr} to authorize.")
    sys.exit(1)   # non-zero exit blocks the tool call

Secondary channel — TTL'd approval file (loud, agent-forgeable — acceptable)

A human can also run an explicit command that writes a stamp:

octo-dim approve-merge 96    # writes ~/.claude/tmp/approved-merge-96.

The hook checks for a fresh stamp (e.g. < 5 min old). This channel is technically agent-forgeable (the agent could call octo-dim approve-merge too), but the canon rule blocks self-approval: the gate script rejects approvals issued by the agent itself. The env channel is the real boundary; the file channel is convenience.

Scoping Rule — Never Blanket, Always Specific

The approved value must encode the exact action:

| Bad (blanket) | Good (scoped) | |---|---| | OCTO_MERGE_APPROVE=yes | OCTO_MERGE_APPROVE=96 | | APPROVE_ALL=1 | OCTO_DEPLOY_APPROVE=prod-2026-06-02 |

A stale OCTO_MERGE_APPROVE=95 cannot authorize PR 96. This prevents approval replay.

Residual Risk

Shell indirection can evade string-based command detection (the part that extracts the PR number from the command):

bash -c "gh pr merge 96"          # hard to parse correctly
eval "gh pr merge 96"             # same

The env channel is immune to this because it doesn't depend on parsing the command — it only checks the env. If the action is truly critical, combine:

  • Env channel for authorization (unforgeable).
  • Command parsing for action identification (parse best-effort; fail-closed on ambiguity).

See [[command-boundary-hook-matching]] for the parsing half.

When to Use

  • Merging PRs that auto-deploy to production.
  • Running destructive DB migrations.
  • Any action where "the agent is doing the work but only the human should pull the trigger."

Reference Implementation

~/.claude/scripts/qa-merge-gate.py — full gate with env channel, file channel, PR-number extraction, and TTL logic.

See also

  • [[command-boundary-hook-matching]] — parsing the command string safely to extract the action identity
  • [[pre-merge-qa-gate]] — QA approval workflow that feeds into this gate
  • [[dry-run-gate-pattern]] — sibling pattern for destructive ops (preview before execute)

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.