AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Ywc Codex Toolkit Eval

skill-yongwoon-ywc-agent-toolkit-ywc-codex-toolkit-eval · by yongwoon

>-

— No reviews yet
0 installs
39 views
0.0% view→install

Install

$ agentstack add skill-yongwoon-ywc-agent-toolkit-ywc-codex-toolkit-eval

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • ✓ Prompt-injection patterns
  • ✓ Secret / credential exfiltration
  • ✓ Dangerous shell & filesystem operations
  • ✓ Untrusted network calls
  • ✓ Known-malicious package signatures

What it can access

  • ✓ Network access No
  • ✓ Filesystem access No
  • ✓ Shell / process execution No
  • ✓ Environment & secrets No
  • ✓ Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-yongwoon-ywc-agent-toolkit-ywc-codex-toolkit-eval)

Reliability & compatibility

✓ Security review passed
0 installs to date
— no reviews yet
● 2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Ywc Codex Toolkit Eval? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

ywc-codex-toolkit-eval

Announce at start: "I'm using the ywc-codex-toolkit-eval skill to evaluate this repository's Codex skills and agents."

This local Codex skill runs a Codex-only evaluate -> improve cycle over codex/skills/* and codex/agents/*.toml. It must stay under .codex/skills/ and must not be copied into codex/skills/ or .codex-plugin/skills/.

Boundary

| Included | Excluded | |---|---| | codex/skills/*/SKILL.md | .claude/skills/* | | codex/skills/*/agents/openai.yaml | claude-code/skills/* | | codex/agents/*.toml | claude-code/agents/* | | Codex install flow, plugin packaging, and bundle docs | Application source code review |

Rationalization Defense

When tempted to bypass a rule, check this table first:

| Excuse | Reality | |---|---| | "scripts/validate.sh is green, so the Codex bundle is healthy" | Validation proves structure and packaging only. Trigger precision, workflow actionability, runtime fit, and behavioral evidence still need graded evaluation. | | "The Claude evaluator already scores Codex paths, so keep using it" | Claude Code and Codex have different frontmatter schemas, agent formats, install paths, and runtime assumptions. Shared concepts are acceptable; shared target roots are not. | | "Local evaluator skills can live under codex/skills if we remember not to publish them" | codex/skills and .codex-plugin/skills are distribution surfaces. Evaluators there will leak into installs or plugin packaging. | | "Mechanical score is enough" | Mechanical axes are partial evidence. Judgment axes such as trigger precision, workflow actionability, scope discipline, mission boundaries, and behavioral evidence still decide the grade. | | "No eval fixtures means no score" | Score from available evidence and mark uncertainty. Missing fixtures are themselves a quality signal. | | "Fix every issue while evaluating" | Evaluation produces an evidence-ranked backlog. Implement fixes through ywc-skill-author, Codex agent authoring, or bundle docs work, then re-score. | | "CI should rewrite history automatically" | CI compares against the reviewed baseline only. Baseline updates are explicit maintenance actions after reviewing the markdown scorecard. |

Violating the letter of these rules is violating the spirit. A quality cycle is useful only when it separates measurement, ownership, and repair.

Arguments

| Parameter | Format | Example | Description | |-----------|--------|---------|-------------| | --target | --target | --target codex/skills | Mechanical scorer target: codex/skills, codex/agents, or all (default). | | --item | --item | --item ywc-plan | Score one Codex skill or agent instead of the whole root. | | --mode | --mode | --mode mechanical | Skill-level evaluation mode: mechanical, judge, or full. Direct score.py runs support mechanical only; judge and full are skill-mediated workflows. | | --ci | flag | --ci | Compare deterministic axes against evals/history.mechanical.json without rewriting it. | | --update-baseline | flag | --update-baseline | Write the reviewed current mechanical baseline after checking the markdown output. | | --only | --only | --only agents | Inventory gate scope: skills, agents, or all. |

Argument Mapping

Map user arguments to concrete commands before running the workflow:

| User request | Inventory command | Mechanical command | |---|---|---| | --target all or no target | inventory_gate.py --json | score.py --mode mechanical --target all --format markdown | | --target codex/skills | inventory_gate.py --json --only skills | score.py --mode mechanical --target codex/skills --format markdown | | --target codex/agents | inventory_gate.py --json --only agents | score.py --mode mechanical --target codex/agents --format markdown | | --item | run the matching inventory scope for context | add --item to score.py and score only that asset in the judgment pass | | --ci | optional unless a full report is requested | score.py --ci | | --update-baseline | optional unless reviewing current structure | score.py --update-baseline after reviewing markdown output |

Direct score.py runs support mechanical mode only. For --mode judge or --mode full, use the mechanical output plus the rubric references to complete the judgment pass.

Scoring Model

Use three layers:

  1. Gate: run scripts/inventory_gate.py to collect inventory and structural pass/fail evidence. Gate failures cap the affected asset at grade C.
  2. Mechanical: run scripts/score.py for deterministic axes and baseline comparison. Judgment axes stay · and final_total remains unavailable.
  3. Judgment: score the remaining axes with the rubric references.

Each dimension is scored 0-4. Weighted composite maps to: A >= 3.5, B 2.5-3.49, C 1.5-2.49, `D

Verdict

  • Status: PASS | PASSWITHACTIONS | REVIEW_REQUIRED | FAIL
  • Assets evaluated:
  • Gate failures:
  • Lowest grade:

Scorecards

| Asset | Kind | Grade | Composite | Weakest dimension | Evidence | |---|---|---:|---:|---|---|

Priority Backlog

  1. [Critical|High|Medium|Low] -

Evidence: Recommended owner:

Scoreboard Update

  • Added:
  • Improved:
  • Regressed:
  • Next review scope:

## Validation

Before claiming the evaluation is complete, verify:

- [ ] `inventory_gate.py` was run for the stated scope.
- [ ] `score.py --mode mechanical --format markdown` was run for mechanical evidence unless skill-level `--mode judge` was explicit.
- [ ] Expected negative checks were run as assertion-shaped commands:
  - `! test -e codex/skills/ywc-codex-toolkit-eval`
  - `! test -e .codex-plugin/skills/ywc-codex-toolkit-eval`
  - `! rg 'tools/codex-internal/skills/ywc-codex-toolkit-[e]val' .codex/skills/ywc-codex-toolkit-eval scripts/validate.sh`
- [ ] Every scored skill used [references/skill-rubric.md](references/skill-rubric.md).
- [ ] Every scored agent used [references/agent-rubric.md](references/agent-rubric.md).
- [ ] Mechanical-only mode is reported as partial, not final quality.
- [ ] Report artifacts, if written, are under `docs/skill-agent-eval/codex/`.
- [ ] No Claude-only paths were scored as Codex assets.
- [ ] `ywc-codex-toolkit-eval` does not appear under `codex/skills/` or `.codex-plugin/skills/`.

## Common Mistakes

- **Evaluating Claude Code artifacts with this skill** — use `.claude/skills/ywc-toolkit-eval` for Claude Code.
- **Treating mechanical points as a final grade** — judgment axes must be scored before claiming a composite grade.
- **Letting local evaluator files leak into packaging** — keep the skill under `.codex/skills/`.
- **Skipping the backlog** — the prioritized fix list is the deliverable that makes the cycle actionable.

## Integration

- **Upstream**: `ywc-skill-author` for Codex skill rule interpretation.
- **Downstream**: use the prioritized backlog to make targeted Codex skill/agent edits, then re-run this skill.
- **Pairs with**: `scripts/validate.sh`, which enforces the local-only packaging boundary and runs the mechanical regression gate.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [yongwoon](https://github.com/yongwoon)
- **Source:** [yongwoon/ywc-agent-toolkit](https://github.com/yongwoon/ywc-agent-toolkit)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.