AgentStack
SKILL verified MIT Self-run

Run Aiewf 2024 Evals

skill-hiteshbandhu-skills-i-use-run-aiewf-2024-evals · by hiteshbandhu

>

No reviews yet
0 installs
2 views
0.0% view→install

Install

$ agentstack add skill-hiteshbandhu-skills-i-use-run-aiewf-2024-evals

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Run Aiewf 2024 Evals? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Run AIEWF 2024 evals & LLM ops

Action playbook from six AI Engineer / World's Fair talks. Do not summarize talks — pick a workflow and execute it.

Supporting files (read when needed):

  • [workflows.md](workflows.md) — workflows A–F (steps, deliverables, stop conditions)
  • [source-index.md](source-index.md) — src-NNN → talk learnings in ingest-into-skills

Optional: {SKILL_OUTPUT_DIR}/run-aiewf-2024-evals/


Step 0 — Pick workflow

Use the decision tree below. Open the matching section in [workflows.md](workflows.md).

What is the user trying to do?
├─ Build domain eval from zero (assertions → judges)     → A
├─ Layer task evals on routers/tools (trace-native)        → B
├─ Calibrate LLM judges / courtroom rubrics              → C
├─ PM+eng eval ops (Zapier/Braintrust regression)          → D
├─ Enterprise CX: full conversation intelligence           → E
└─ Decide fine-tune vs prompt/distill (maturity curve)     → F

Stop summarizing once a workflow is identified — run its checklist.


Install

cp -r skills/run-aiewf-2024-evals ~/.claude/skills/
cp -r skills/run-aiewf-2024-evals ~/.cursor/skills/
cp -r skills/run-aiewf-2024-evals ~/.codex/skills/

From skills-i-use or ingest-into-skills (playlists/evals-llm-ops-aie-world-s-fair-2024/).


Cross-cutting rules

| Rule | Source | |------|--------| | Evals ≠ demos; log before judges | [src-001 @ 6:39] | | Assertions before LLM-as-judge | [src-001 @ 4:41] | | Layer scores on router/tools/answer | [src-004 @ 4:46] | | Fine-tune only after eval + teacher data | [src-006 @ 5:55] |

Disputed steps: see [source-index.md](source-index.md). Name workflow A–F; save artifacts to ./skill-outputs/run-aiewf-2024-evals/ when requested; do not auto-commit.


Invocation examples

@run-aiewf-2024-evals build domain eval harness for our agent
should we fine-tune or stay on GPT-4o?

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.