AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Best Version Builder

skill-cyfung1031-skills-v1-3-2 · by cyfung1031

Compare, score, synthesize, retest, and plateau-stop candidate versions with source-grounded identity, target maps, calibrated objectives, broad category buckets with explicit coverage lenses, validation-tier claim caps, evidence ledgers, recomputable score packets, observability markers, artifact checksums, correction audits, and anti-compression safeguards.

No reviews yet
0 installs
14 views
0.0% view→install

Install

$ agentstack add skill-cyfung1031-skills-v1-3-2

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-cyfung1031-skills-v1-3-2)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
22d ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Best Version Builder? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Candidate Version Analysis + Plateau Synthesis

Use when the user provides candidate versions, prompts, skills, specs, code variants, documents, archives, benchmark reports, prior matrices, or generated successors and asks to compare, score, choose, improve, synthesize, retest, correct, or validate a best version.

Core invariant: identity → target map → objective → broad rubric → coverage lenses → evidence → measurement → observability → recomputation → synthesis → correction → handoff. Keep these gates separate. Do not score from filenames, recency, polish, length, prior summaries, or assumed intent. Do not rewrite first and rationalize afterward.

1. Intake, Identity, and Workspace

  1. Use the requested workspace; otherwise use the active project workspace. Unpack archives into a clean source directory. Preserve originals unchanged; write generated successors and results to separate output paths.
  2. Ignore noise: __MACOSX, .DS_Store, resource forks, temp files, caches, build products, duplicate outputs, and prior generated artifacts unless explicitly in scope.
  3. Inventory every candidate: name/version/path, checksum or stable size marker, artifact type, declared purpose, comparison unit, output contract, protected tokens, constraints, estimator/estimate, scope notes, category-lens coverage notes, observability hooks, and source-vs-generated boundary.
  4. Identify candidates from source content, not filename alone. If filename, frontmatter, title, and body disagree, report the conflict and use content-grounded identity.
  5. Mark each artifact in-scope, comparator-only, or out-of-scope. Comparator-only artifacts may inform analysis but cannot silently win for another target.
  6. Before final handoff, verify every saved path exists and report checksum/status. Never invent sandbox links or silently overwrite source or prior generated files.

2. Target Behavior Map

Before scoring, write a compact target map:

  • intended user task, artifact purpose, comparison unit, activation/stop conditions, and non-goals;
  • required outputs, saved artifacts, handoff shape, protected tokens, exact-preservation needs, and localization expectations;
  • safety/professional, correctness, evidence, validation, observability, audit, reproducibility, compatibility, and correction requirements;
  • critical legacy behaviors that must survive synthesis;
  • optional helpful behaviors;
  • irrelevant, comparator-only, and wrong-target features that must not be rewarded or penalized unless explicitly required.

Rules: no target map caps production recommendation at 8.0. Unresolved target mismatch caps production at 6.5. A prior matrix using another target map is critique evidence only. Score target fit separately from feature richness. Mark every category and material feature as required, optional, conditional, out-of-scope, or comparator-only.

3. Objective, Broad Rubric, Coverage Lenses, Weights, Formula

Default objective: best production version.

Profiles:

  • production: reliable target-task performance, correctness, evidence, observability, auditability, maintainability, compatibility, correction quality; cost is secondary.
  • compressed: cost minimization only after required production behavior remains acceptable.
  • research/exploratory: diagnostic coverage and alternatives; higher cost tolerance.
  • live-model: repeated model/API calls with variance notes.
  • deterministic proxy: static/heuristic checks only; never imply live behavior.

Default broad categories when the user gives none. Keep the top-level rubric broad enough to transfer across benchmark styles, artifacts, and domains; use coverage lenses to prevent blind spots instead of multiplying narrow mandatory categories.

| Category | Weight | Production-critical | Coverage lenses to inspect | |---|---:|---|---| | Correctness / target behavior | 1.25 | Yes | Executability, semantic preservation, required outputs, exact content, constraints, edge cases, stop conditions, contradiction checks. | | Scope / user value / fit | 1.15 | Yes | Target-map fit, comparison-unit fit, non-goals, comparator-only handling, user value, wrong-target expansion. | | Safety / guardrails / risk control | 1.15 | Yes | Safety, privacy, professional-risk boundaries, protected tokens, overwrite risk, false certainty, abuse resistance. | | Evidence / validation / reproducibility | 1.20 | Yes | Source grounding, validation tier, fixtures, pass/fail criteria, estimator stability, formulas, seeds, versions, checksums, recompute ability. | | Observability / audit / handoff | 1.05 | Yes | Failure localization, logs, metrics, evidence rows, artifact paths, manifests, score packets, checksums, correction notes, claim limits. | | Maintainability / evolvability | 1.00 | Yes | Modularity, readability, low duplication, update safety, cognitive load, anti-overfitting, reusable gates. | | Compatibility / regression safety | 1.00 | Yes | Legacy behavior, output contracts, localization, exact-preservation rules, migration/rollback, regression firewall. | | Cost / complexity / token efficiency | 0.65 | No unless compressed objective | Token use, runtime, artifact burden, maintenance burden, harness complexity, simplicity after required behavior is safe. | | Language / audience, when relevant | 1.00 | Conditional | Requested language, terminology, domain terms, reader level, localization, code/path/identifier preservation. |

Coverage-lens rule: do not create a new top-level category just because a lens matters. First decide whether the lens is required, optional, conditional, out-of-scope, or comparator-only under the target map; then score it inside the broad category where it best affects the outcome. Add a narrow category only when the user asks for it, the target map requires separate reporting, or merging it would materially hide a winner-changing defect.

Category-map rule: before scoring, mark each broad category and material lens required, optional, conditional, out-of-scope, or comparator-only; define category/lens-specific evidence requirements and caps. If old and new matrices use different categories, weights, anchors, relevance labels, or pass/fail criteria, either backfill old candidates under the new rubric or label the comparison non-equivalent. Do not let token efficiency, polish, or feature richness compensate for a production-critical category below the declared pass threshold.

Default 10/8/6/4 anchors unless the user provides better anchors:

  • 10: complete target fit; strong positive source or measured evidence; no material missing requirement; observable, reproducible, auditable, and regression-safe.
  • 8: production-usable with minor gaps; evidence supports the main claim; limitations are disclosed and do not threaten required behavior.
  • 6: partially useful but missing a material requirement, validation, observability, reproducibility, or compatibility evidence; safe only with caveats.
  • 4: weak, ambiguous, wrong-target, hard to execute, or mostly unsupported; likely to fail important cases.

Use user categories exactly when provided, but warn when they omit production-critical lenses such as correctness, observability, reproducibility, or compatibility. When user categories are narrow, keep them for reporting and add a separate broad-category rollup if it improves comparability. Before scoring, state relevance, weights, anchors, pass/fail criteria, estimator, formula, fixture checksum, validation tier, and limits. Prefer tokenizer counts; otherwise use a stable word-count proxy and label it. Do not change estimator mid-run. For compressed objectives, increase token weight only after all production-critical categories are at least 7.5.

Default formula when no better user formula exists:

combined = production_effectiveness - cost_penalty + coverage_bonus - regression_penalty

Define each term once. Store raw data so another run can recompute the rank. Production effectiveness should be the weighted mean of required and conditional production-critical categories, not a raw feature count.

4. Validation Tier and Claim Caps

Classify before scoring:

| Tier | Evidence | Allowed claim | Cap | |---|---|---|---:| | T0 | source inspection only | best structural/proxy candidate | 9.2 | | T1 | fixed fixtures + deterministic/heuristic scoring | best deterministic-proxy measured candidate | 9.4 | | T2 | repeated live model/API calls + variance notes | best live-measured candidate for this setup | 9.7 | | T3 | qualified independent human review | best validated under review protocol | none |

Separate production effectiveness, overall including cost, confidence, and tier. Never imply live behavior from T0/T1. If the winner changes under another tier, estimator, objective, weights, target map, category/lens map, or fixture coverage, report sensitivity. A challenged result must show old claim, corrected claim, and obsolete conclusion.

5. Evidence, Caps, and Score Packet

Score 0-10. Scores above 8 require positive source evidence; above 9 require strong evidence and no material missing requirement. If a category cannot be evaluated, score conservatively and label the limit.

Caps unless stronger source evidence justifies override:

  • full comparison unit not read: max overall 6.0;
  • unresolved identity/scope mismatch: max production 6.5;
  • no/wrong target map: max production 8.0;
  • undefined category relevance in a disputed score: max affected category 7.5;
  • missing correctness evidence for a production recommendation: max correctness 7.0 and max production 8.0;
  • no observability path for failures in recommendation-grade work: max observability 6.5 and max production 8.5;
  • no evidence ledger for recommendation-grade work: max evidence 6.5 and audit 7.0;
  • no manifest for large/repeated work: max handoff/audit 7.0;
  • different inputs, formulas, estimators, weights, category/lens maps, target maps, or relevance labels: no best-measured claim;
  • token efficiency is the only winning production category: label best compressed, not best production;
  • deterministic proxy presented as live behavior: invalidate until corrected.

For recommendation-grade, repeated, challenged, synthesized, or large benchmark work, save enough to audit and recompute:

  • run_manifest.json: run id, timestamp, source/generated paths, ignored noise, checksums/sizes, objective, target map, category relevance, anchors, weights, formula, estimator, caps, validation tier, fixture checksum/count, seed, model/API settings if any, validation commands, observability markers, failure taxonomy, exclusions, overwrite policy, old-vs-corrected references.
  • candidate_inventory.csv: identity, path, checksum, purpose, comparison unit, token estimate, target-fit status, category/lens coverage status, observability hooks.
  • category_map.json: category names, weights, relevance labels, production-critical flags, anchors, pass/fail criteria, evidence requirements, cap rules, and any user overrides.
  • evidence_ledger.csv: candidate, category, relevance, score, source locator, positive evidence, missing evidence, confidence, evidence type, validation tier.
  • score_matrix.csv: category scores, production effectiveness, cost terms, overall/combined, pass rate, failure count, rank.
  • row_results.csv/jsonl when fixtures exist: candidate, fixture id/category, input size, target behavior, expected behavior, score, cost, failures, failure locality, debug evidence, notes, evidence marker, validation tier.
  • iteration_log.csv/json: revision target, old/new scores, deltas, changed sections, expected benefit, regressions, accept/reject reason, rollback path, probe, stop reason.
  • correction_audit.csv/json for challenged scores.

When feasible, include a small rank-recompute script or command. Run it once. Prose, matrix, ledger, category/lens map, manifest, observability markers, and recomputed rank must agree.

6. Analysis Gate

For each candidate, read the complete comparison unit and report: intent, mechanism, constraints, output shape, validation behavior, assumptions, protected tokens, source-vs-generated boundary, target-map fit, category/lens coverage, strengths, weaknesses, best use case, likely failure mode, observability/debug path, cost, comparability, objective fit, uncertainty, and score rationale.

Penalize wrong-target scoring, hidden ambiguity, unsupported claims, duplicated/conflicting rules, stale text, untestable guidance, missing correctness behavior, missing safety behavior, weak observability, weak handoff, over-compression, unverifiable superiority claims, localization loss, source/result overwrite risk, unrecomputable rank, and challenge-handling gaps. Reward executable sequencing, target-fit clarity, priority ladders, exact code/error/log/token preservation, domain-term preservation, validation hooks, evidence-grounded scoring, auditability, failure traceability, compactness without loss, recomputable packets, and legacy compatibility. Do not bias toward newer, longer, shorter, or more polished versions.

Prior reports, user challenges, and old matrices are critique evidence, not ground truth.

7. Measurement Gate

Use when the user asks for best, best measured, test again, same mixture, complete matrix, effectiveness, token consumption, repeated improvement, synthesis validation, or a successor.

  1. Reuse exact prior fixtures when requested; verify checksum.
  2. Otherwise build fixtures from the target map and category/lens map. Use the requested count; if none is given, use a compact diverse set sufficient for every active category.
  3. Include normal, edge, adversarial, multilingual, fixed-format, exact code/error/log, domain terms, risk/safety, ambiguity, protected tokens, sourced claims, audit/file handoff, rewrite, summary, decision, procedure, localization, observability, reproducibility, compatibility, and legacy-regression cases as relevant.
  4. Measure candidates and successors with identical inputs, target map, category/lens map, relevance labels, weights, formula, estimator, pass/fail criteria, validation tier, and failure taxonomy.
  5. Track effectiveness and cost separately: skill tokens, prompt/output tokens, runtime, file size, maintenance burden, artifact count, and rerun complexity as relevant.
  6. Track observability separately: failure taxonomy, error locality, debug evidence, log/row locator, artifact status, and whether another reviewer can reproduce the failure from the packet.
  7. Label deterministic proxy, heuristic, live-model/API, or human-review status clearly.
  8. Re-rank after adding any candidate or successor. Recommend the best measured version under the declared objective and target map, not the newest or shortest.

8. Sanity, Sensitivity, and Dominance

Before publishing or synthesizing:

  1. Recompute means, weighted scores, costs, bonuses, penalties, pass rates, and combined scores from raw rows.
  2. Confirm identical inputs, categories, weights, formula, estimator, pass/fail rules, fixtures, target map, relevance labels, candidate inclusion rules, observability markers, and validation tier.
  3. Spot-check evidence for top two, bottom two, newly added candidates, close-margin winners, and material rank changes.
  4. Compare rank with target map and dominance rules. Resolve wrong-target, compression-only, estimator-sensitive, low-evidence, low-observability, and validation-tier-sensitive wins.
  5. Run sensitivity checks for close margins by varying token-cost and subjective weights within a declared range; call the result margin-sensitive when ranks change.
  6. Check contradiction triad: prose claims, matrix rows, category/lens map, manifest, and evidence ledger must agree with objective, target map, estimator, and validation tier.
  7. Do not mix token counts, word proxies, and file sizes as equivalent.

9. Synthesis and Regression Firewall

Synthesize only after analysis/measurement evidence exists. Pres

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.