Install
$ agentstack add skill-tikalk-adlc-team-skills-evals-validate ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
evals-validate
What this skill does
Conducts comprehensive validation of the implemented evaluation system following EDD principles to ensure production readiness through statistical analysis, performance verification, and quality assurance.
Output:
- Statistical Validation - TPR/TNR analysis, accuracy metrics, confidence intervals
- Performance Validation - SLA compliance verification for evaluation pyramid tiers
- Quality Assurance - Goldset integrity, example balance, coverage analysis
- Holdout Dataset Validation - Unbiased accuracy assessment on reserved test set
- Auto-handoff to
/evals-analyzefor closed loop trajectory analysis
Key EDD Principles Applied:
- Principle IV: Evaluation Pyramid - Tier performance SLA validation (Tier 1 <30s, Tier 2 <5min)
- Principle II: Binary Pass/Fail - Statistical compliance verification
- Principle IX: Test Data as Code - Holdout dataset validation integrity
- Principle III: Error Analysis - Pattern stability validation
When to use
- After
/evals-implement: Execute the evaluation suite and measure quality - CI/CD Pipeline gate: Run evaluations before release to ensure no regressions
- Periodic audit: Verify evaluator accuracy on holdout data to check for model drift
When NOT to use
- Evaluator not generated: Run
/evals-implementto build grader files first - Analysing failure traces: Use
/evals-analyzeto extract deep insights from run results
Process
User Input
$ARGUMENTS
--holdout-only— Validate only on holdout dataset (unbiased validation)--performance-only— Skip statistical analysis, focus on SLA compliance--metrics METRICS— Specific metrics to validate (tpr, tnr, accuracy, performance)
Execution Steps
Phase 1: Execute Evaluations
Runs the underlying framework CLI directly:
- PromptFoo:
npx promptfoo eval --config evals/promptfoo/config.js - DeepEval:
pytest evals/deepeval/ -vorpython evals/deepeval/config.py
Phase 2: Compute Statistical Validation
- Parse generated results JSON from
evals/results/. - Calculate True Positive Rate (TPR) and True Negative Rate (TNR).
- Calculate overall accuracy with 95% confidence intervals.
- Ensure no Likert scales or numerical scores leak into results.
Phase 3: SLA Compliance Check
- Measure execution times for Tier 1 and Tier 2.
- Verify Tier 1 completes under 30 seconds.
- Verify Tier 2 completes under 5 minutes.
- Check headroom analysis (SLA budget consumed).
Phase 4: Write Validation Report
- Write validation results to
evals/results/validation_report.md. - Include pass/fail counts, TPR/TNR table, SLA timings, and holdout set results.
Phase 5: Auto-Handoff
Trigger /evals-analyze to close the loop.
Verification
- Evaluation execution successfully completed with results JSON written to
evals/results/ evals/results/validation_report.mdcreated with TPR/TNR and SLA metrics- Statistical metrics calculated with confidence intervals
- Headroom and SLA compliance verified
- Handover summary lists results and validation report path
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: tikalk
- Source: tikalk/adlc-team-skills
- License: MIT
- Homepage: https://github.com/tikalk/agentic-sdlc-12-factors
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.