Install
$ agentstack add skill-skillberry-ai-cap-evolve-gate ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
gate — accept only real improvements, on val
The gate is where dishonest optimization is prevented. Search is a noise amplifier: try enough candidates and some will look better by chance alone (the more candidates you screen, the larger the expected best-of-noise). The gate is the rule that keeps a lucky draw from being promoted to "the new best". It refuses any split but val, and by default accepts a candidate only when its val reward beats the current best by more than k standard errors.
Inputs / outputs (manifest tokens)
- needs:
scores— the candidate's and current best's val reward and
stderr (from evaluate). The SE is not optional: significance is meaningless without it.
- provides:
decision—{accept, reason, delta, threshold}, the audit
record of why a candidate was kept or rejected.
The significance rule (default)
SE = sqrt(candidate_stderr^2 + current_stderr^2) # SE of the difference
accept ⟺ Δ = candidate_val − current_val > k · SE
This is the standard test for "is the difference real?": the SE of a difference of two independent means is the root-sum-of-squares of their SEs, and Δ > k·SE asks whether the gap clears k standard errors of that difference. k=1 is lenient (~1σ); raise it to be stricter. It is the textual-optimization analogue of Koehn's bootstrap significance test for metric differences — accept only when the gap is unlikely to be noise.
Single-trial scores report stderr=0, collapsing k·SE to 0 — then significant silently degrades to strict and accepts any positive blip. If you run the significance gate, score with multiple trials (see evaluate).
Modes
significant(default):Δ > k·SE— variance-aware, the honest choice.strict:Δ > 0— any improvement. Only safe with a near-zero-variance scorer
(deterministic, single correct answer).
threshold:Δ > T— a flat margin (use when you have a domain minimum
worthwhile gain, e.g. "don't bother unless +2pp").
simplicity_tiebreak: like strict, but on a (near-)tie prefer the smaller
candidate — an Occam bias against bloated edits that don't earn their size.
No-regression (the second gate)
A mean can rise while previously-passing tasks silently break. Pair the significance gate with a no-regression check: reject a candidate that improves the aggregate but drops any task that the current best passed. This is the same dual-gate discipline SWE-bench-style harnesses use (a patch must pass the new tests and not break the existing ones — FAILTOPASS and PASSTOPASS). diagnose provides kept_good (the currently-passing tasks) precisely so this check has something to protect.
Dual-mode
This phase runs two ways from the same SKILL.md: standalone as the slash command /cap-evolve:gate (the argument-hint shows its run.py args), and orchestrator-callable — cap-evolve run / the orchestrate skill invokes the same scripts/run.py headlessly and threads the run dir between phases.
How to run
python scripts/run.py --current 0.50 --candidate 0.62 \
--mode significant --k-se 1.0 --candidate-stderr 0.03 --current-stderr 0.03
Algorithms call the gate internally every iteration via the harness; this skill exists so a human or agent can reproduce and understand a single decision.
What good vs bad looks like
- Good:
significantmode with real multi-trial SEs; a no-regression check on
top; every accept/reject logged with its reason.
- Bad: gating on
train(the tool refuses this — it overfits the optimizer to
the data it edits against); strict mode on a noisy agent (accepts noise); raising the mean while quietly regressing tasks because no-regression was off.
References
references/concepts.md— the difference-of-means SE, choosingk, the
multiple-comparisons motivation, the dual-gate / no-regression rationale, and why gating on val (never train, never test) is the honest split, with sources.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: skillberry-ai
- Source: skillberry-ai/cap-evolve
- License: Apache-2.0
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.