# Evals Init

> Initialize evals/{system}/ directory structure for evaluation system following EDD principles (Standalone). Choose PromptFoo or DeepEval based on tech stack, generate security baseline.

- **Type:** Skill
- **Install:** `agentstack add skill-tikalk-adlc-team-skills-evals-init`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [tikalk](https://agentstack.voostack.com/s/tikalk)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [tikalk](https://github.com/tikalk)
- **Source:** https://github.com/tikalk/adlc-team-skills/tree/main/skills/evals/evals-init
- **Website:** https://github.com/tikalk/agentic-sdlc-12-factors

## Install

```sh
agentstack add skill-tikalk-adlc-team-skills-evals-init
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# evals-init

## What this skill does

Initialize the **project-level evaluation directory structure** following EDD (Eval-Driven Development) principles to prepare for systematic evaluation development. This is completely standalone with zero spec-kit dependencies.

**Output**:
1. **Directory Structure** - `evals/{system}/` with proper organization (promptfoo | deepeval)
2. **Security Baseline** - Auto-created graders for PII leakage, prompt injection, hallucination detection, misinformation detection
3. **Configuration Files** - Standalone config.yml and goldset templates under `.adlc/evals/`
4. **Auto-handoff** to `/evals-specify` to begin error analysis

**Key EDD Principles Applied**:
- **Principle I**: Spec-Driven Contracts - Evals validate spec compliance
- **Principle II**: Binary Pass/Fail - No Likert scales in grader templates
- **Principle IV**: Evaluation Pyramid - Tier 1 (fast) + Tier 2 (goldset) structure
- **Principle IX**: Test Data as Code - Version control setup for datasets

## When to use

- **Starting systematic evaluation**: Set up the initial evaluation harness for your application
- **EDD Adoption**: Converting from traditional testing to evaluation-driven development
- **Security-first evaluation**: Auto-generate baseline security checks from the start

## When NOT to use

- **Evals directory already exists**: Use `/evals-validate` to run tests, or `/evals-specify` to add criteria
- **Evaluating team directives**: This is for project-level application behavior testing, not directives compliance

## Process

### User Input
```text
$ARGUMENTS
```
Parse flags from the arguments first, then treat remaining text as focus areas:
- `--system SYSTEM` — Choose `promptfoo` or `deepeval`. If omitted, choose interactively based on tech stack.
- Remaining text — System description (focus setup)

### Execution Steps

#### Phase 1: Tech Stack Detection
- Scan project manifests (`package.json`, `requirements.txt`, `Cargo.toml`, `go.mod`, etc.)
- Recommends PromptFoo for mixed/JS stacks; DeepEval for Python-native stacks

#### Phase 2: Create Directory Structure
Creates:
```
evals/
├── {system}/                    # promptfoo | deepeval
│   ├── goldset.md              # Published goldset
│   ├── goldset.json            # Auto-generated for system consumption
│   ├── config.yml              # System-specific configuration
│   ├── config.{js,py}          # Generated system config (.js for promptfoo, .py for deepeval)
│   └── graders/                # Binary pass/fail graders
│       ├── check_pii_leakage.py           # Security baseline
│       ├── check_prompt_injection.py     # Security baseline
│       ├── check_hallucination.py        # Security baseline
│       └── check_misinformation.py       # Security baseline
├── results/                    # Git-ignored run outputs
└── .adlc/
    └── drafts/evals/           # Draft eval records (Markdown + YAML)
```

#### Phase 3: Configuration Copy
- Create `.adlc/evals/` if missing.
- Copy `skills/evals/evals-templates/evals-config-template.yml` to `.adlc/evals/evals-config.yml`.

#### Phase 4: Auto-Handoff
Trigger `/evals-specify` to begin error analysis.

## Verification
- `evals/{system}/goldset.md` exists (initially empty)
- `.adlc/evals/evals-config.yml` exists
- Graders directory populated with 4 security baseline python scripts
- Results directory contains `.gitignore` to prevent versioning traces
- Handover report generated with recommended framework and next steps

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [tikalk](https://github.com/tikalk)
- **Source:** [tikalk/adlc-team-skills](https://github.com/tikalk/adlc-team-skills)
- **License:** MIT
- **Homepage:** https://github.com/tikalk/agentic-sdlc-12-factors

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-tikalk-adlc-team-skills-evals-init
- Seller: https://agentstack.voostack.com/s/tikalk
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
