AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Auto Skill Build Agent Reliability Loop

skill-arnie016-codex-prompt-templates-auto-skill-build-agent-reliability-loop · by Arnie016

>

No reviews yet
0 installs
35 views
0.0% view→install

Install

$ agentstack add skill-arnie016-codex-prompt-templates-auto-skill-build-agent-reliability-loop

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-arnie016-codex-prompt-templates-auto-skill-build-agent-reliability-loop)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
4mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Auto Skill Build Agent Reliability Loop? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Agent Reliability Loop

Generated by: Codex Supercharge maintenance automation.

Goal: turn an agent workflow into a closed loop where production traces become eval cases, eval failures become guardrails or fixes, and gateway policy keeps cost, routing, and tool risk bounded.

Skip When

  • The user only wants local Codex run/cost tracking.
  • The work is a single benchmark or experiment with no production surface.
  • The task is only MCP conformance; use $auto-skill-build-mcp-conformance-harness.

Workflow

  1. Map the surface: list user journeys, models, tools, data classes,

side effects, risk levels, and quality/cost/latency targets.

  1. Instrument first: require request IDs, trace/span IDs, session IDs,

model/provider, token/cost, cache, fallback, guardrail, and tool metadata.

  1. Build eval bundles: create golden cases for task success, prompt

conformance, tool correctness, unsafe requests, PII/secrets, latency, and cost. Keep judge prompts and heuristic checks versioned.

  1. Set gateway policy: define provider mapping, model fallbacks, cache

rules, virtual keys, budgets, rate limits, privacy redaction, and audit logs.

  1. Stage guardrails: run new pre/post checks in log or monitor mode first,

then enforce only after false positives and fail-open/fail-closed behavior are explicit.

  1. Shadow safely: mirror sampled traffic to candidate models or prompts

only when shadow calls cannot trigger external side effects.

  1. Close the loop: cluster failed traces, match nearest successful traces,

add representative failures to evals, patch prompts/tools/policies, then rerun the bundle before release.

Commands

rg -n "trace|span|cost|tokens|guardrail|fallback|budget|rate limit|request_id" .
rg -n "eval|rubric|judge|golden|dataset|experiment|shadow|canary" .
rg -n "tool|mcp|side effect|webhook|shell|filesystem|credential" .

Output

# Agent Reliability Plan
## Surface
## Instrumentation
## Eval Bundle
## Gateway Policy
## Guardrail Rollout
## Shadow Or Canary Plan
## Trace-To-Regression Loop
## Risks And Trust Notes
## Validation

Validation

  • Every production route has at least one trace and one regression case.
  • Every blocking guardrail has a false-positive review path.
  • Every budget/rate limit has an owner and an alert threshold.
  • Shadow/canary traffic cannot write to tools, accounts, payments, or user data.

Read references/future-agi-agent-reliability-loop.md for the source-backed pattern and risk notes.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.