AgentStack
SKILL verified MIT Self-run

Auto Skill Build Agent Reliability Loop

skill-arnie016-codex-prompt-templates-auto-skill-build-agent-reliability-loop · by Arnie016

>

No reviews yet
0 installs
3 views
0.0% view→install

Install

$ agentstack add skill-arnie016-codex-prompt-templates-auto-skill-build-agent-reliability-loop

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Auto Skill Build Agent Reliability Loop? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Agent Reliability Loop

Generated by: Codex Supercharge maintenance automation.

Goal: turn an agent workflow into a closed loop where production traces become eval cases, eval failures become guardrails or fixes, and gateway policy keeps cost, routing, and tool risk bounded.

Skip When

  • The user only wants local Codex run/cost tracking.
  • The work is a single benchmark or experiment with no production surface.
  • The task is only MCP conformance; use $auto-skill-build-mcp-conformance-harness.

Workflow

  1. Map the surface: list user journeys, models, tools, data classes,

side effects, risk levels, and quality/cost/latency targets.

  1. Instrument first: require request IDs, trace/span IDs, session IDs,

model/provider, token/cost, cache, fallback, guardrail, and tool metadata.

  1. Build eval bundles: create golden cases for task success, prompt

conformance, tool correctness, unsafe requests, PII/secrets, latency, and cost. Keep judge prompts and heuristic checks versioned.

  1. Set gateway policy: define provider mapping, model fallbacks, cache

rules, virtual keys, budgets, rate limits, privacy redaction, and audit logs.

  1. Stage guardrails: run new pre/post checks in log or monitor mode first,

then enforce only after false positives and fail-open/fail-closed behavior are explicit.

  1. Shadow safely: mirror sampled traffic to candidate models or prompts

only when shadow calls cannot trigger external side effects.

  1. Close the loop: cluster failed traces, match nearest successful traces,

add representative failures to evals, patch prompts/tools/policies, then rerun the bundle before release.

Commands

rg -n "trace|span|cost|tokens|guardrail|fallback|budget|rate limit|request_id" .
rg -n "eval|rubric|judge|golden|dataset|experiment|shadow|canary" .
rg -n "tool|mcp|side effect|webhook|shell|filesystem|credential" .

Output

# Agent Reliability Plan
## Surface
## Instrumentation
## Eval Bundle
## Gateway Policy
## Guardrail Rollout
## Shadow Or Canary Plan
## Trace-To-Regression Loop
## Risks And Trust Notes
## Validation

Validation

  • Every production route has at least one trace and one regression case.
  • Every blocking guardrail has a false-positive review path.
  • Every budget/rate limit has an owner and an alert threshold.
  • Shadow/canary traffic cannot write to tools, accounts, payments, or user data.

Read references/future-agi-agent-reliability-loop.md for the source-backed pattern and risk notes.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.