# Context Benchmark

> Measure context preparation across fixed cases. Use when token savings, hard-budget compliance, protected-fact recall, evidence anchors, determinism, runtime, or downstream answer quality must be demonstrated with reproducible evidence rather than marketing claims.

- **Type:** Skill
- **Install:** `agentstack add skill-tikazi-tikaz-ai-skills-context-benchmark`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [TIKAZI](https://agentstack.voostack.com/s/tikazi)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [TIKAZI](https://github.com/TIKAZI)
- **Source:** https://github.com/TIKAZI/TIKAZ-AI-Skills/tree/main/suites/context-economy/context-benchmark
- **Website:** https://tikazi.github.io/TIKAZ-AI-Skills/

## Install

```sh
agentstack add skill-tikazi-tikaz-ai-skills-context-benchmark
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Context Benchmark

Designed, integrated, independently refactored, and continuously maintained by **TIKAZ**.

Run a versioned manifest of independent cases and keep raw per-case results. Report efficiency and quality separately:

- source and packed tokens;
- final-budget compliance;
- protected-fact recall;
- evidence-anchor correctness;
- deterministic repeatability;
- preparation runtime;
- optional externally supplied answer score.
- document-route correctness, informative-visual recall, decorative/duplicate skip accuracy, and complex-table fidelity warnings for multimodal fixtures.

Do not hide failures inside averages. A smaller pack with lower fidelity is a regression, not a win. Do not claim superiority until the same files, questions, model/detail settings, budgets, and blind answer rubric are used. Use the shared CLI `benchmark` command and read `../references/benchmark-method.md` when publishing results.

Publish `benchmarks/results/metrics.json`, the generated evidence card, and raw cases together. Keep context efficiency, exact-repeat prompt efficiency, literal fact/anchor fidelity, multimodal routing, and pending provider/vision/downstream evidence separate; never replace them with one composite fidelity score.

From the suite directory, run `python scripts/tikaz_context.py benchmark --manifest benchmarks/manifest.json --output `. Inspect both `summary.json` and `cases.json`; a passing case is not evidence of positive savings or semantic equivalence.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [TIKAZI](https://github.com/TIKAZI)
- **Source:** [TIKAZI/TIKAZ-AI-Skills](https://github.com/TIKAZI/TIKAZ-AI-Skills)
- **License:** MIT
- **Homepage:** https://tikazi.github.io/TIKAZ-AI-Skills/

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-tikazi-tikaz-ai-skills-context-benchmark
- Seller: https://agentstack.voostack.com/s/tikazi
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
