AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified Apache-2.0 Self-run

Codex Session Benchmark Maintainer

skill-undertone0809-rudder-codex-session-benchmark-maintainer · by Undertone0809

Use when benchmarking local Codex sessions against recent history: target session comparisons, recent 30/50/100 session cohorts, efficiency, follow-up rate, interruption, rework, problem-resolution proxies, workflow quality, or skill/workflow improvement signals.

No reviews yet
0 installs
6 views
0.0% view→install

Install

$ agentstack add skill-undertone0809-rudder-codex-session-benchmark-maintainer

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-undertone0809-rudder-codex-session-benchmark-maintainer)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Codex Session Benchmark Maintainer? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Codex Session Benchmark Maintainer

Overview

Compare a target Codex session or session class against a clean local cohort using explicit proxy metrics and caveats.

When to Use

Use this skill when:

  • the user gives a Codex session id and asks how it compares
  • the user asks for recent 30/50/100 session benchmark or efficiency analysis
  • the desired output is proxy metrics, failure classes, and workflow improvement recommendations
  • session quality must be compared against local history rather than judged in isolation

Do not use this skill when:

  • review-only product verdicts; use codex-session-product-reviewer-maintainer
  • cohort-only skill hygiene when the deliverable is which skill to optimize
  • raw conversation search with no benchmark question

Core Pattern

target -> clean cohort -> proxy metrics -> baseline comparison -> workflow recommendation

Quick Reference

| Situation | Action | | --- | --- | | Target session supplied | Resolve logs and session metadata first | | Recent cohort requested | Deduplicate roots and exclude current/spawned noise | | Outcome metric requested | Label proxy metrics and caveats | | Skill recommendation requested | Map repeated failures to existing owners before proposing new skills |

Implementation

  1. Resolve the target session and local rollout path.
  2. Build a clean comparable cohort from local SQLite/session evidence.
  3. Extract proxy metrics such as turns, follow-ups, interruptions, spawned children, commits, and validation evidence.
  4. Classify outcome and failure modes with caveats.
  5. Recommend the smallest workflow or skill improvement supported by the cohort.

Reference files are part of this skill contract. Before executing high-risk actions or final judgments, load references/runbook.md for the detailed legacy workflow, examples, validation cases, and command-level guidance.

Use evals/ when the route needs that detail; keep the entrypoint thin.

Common Mistakes

| Mistake | Fix | | --- | --- | | Treating proxy metrics as true satisfaction | Name the caveat and avoid overclaiming. | | Letting spawned children pollute the cohort | Collapse to root sessions first. | | Benchmarking without a target or cohort boundary | Ask or infer the smallest defensible cohort. | | Inventing a new skill from one cluster | Compare against existing maintainers first. |

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.