# Jev Harness

> Zero-dependency System One decision harness: 5 semantic gates saving frontier AI agent tokens on trivial errors & doom loops. Python + TypeScript + Rust. MCP-compatible.

- **Type:** MCP server
- **Install:** `agentstack add mcp-ismaelsoilet-jev-harness`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [ismaelsoilet](https://agentstack.voostack.com/s/ismaelsoilet)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [ismaelsoilet](https://github.com/ismaelsoilet)
- **Source:** https://github.com/ismaelsoilet/jev-harness
- **Website:** https://pypi.org/project/jev-harness/

## Install

```sh
agentstack add mcp-ismaelsoilet-jev-harness
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# ⚡ Jev Harness: The Token Optimizer & Decision Gate for AI Coding Agents

  
  
  
  
  
  
  
  
  
  
  
  
  
  

  🇬🇧 English | 🇧🇷 Português

> **Stop burning 50,000 frontier tokens on missing packages, network flakiness, and circular doom loops.**
>
> `jev-harness` is an ultra-fast, deterministic-first token optimizer, test failure triage gate, and semantic guardrail for AI coding agents (OpenCode, Command Code, Claude Code, Cursor, Antigravity IDE, Windsurf, Zed, and Pi). Combines instant zero-token offline deterministic heuristics (
> 🤖 **Operating as an AI Agent?** If you are **working on this repository**, read [`AGENTS.md`](AGENTS.md) ([Português](AGENTS.pt-BR.md)) first. If you are **integrating the tool into your own project**, you do not need it — use the [**Universal AI Agent Integration Guide**](docs/AGENT_INTEGRATION_GUIDE.md) (2-minute setup, no API key required to start).
>
> **What it is:** a deterministic-first **System 1.5 decision layer** for the code-quality loop — test-failure triage, doom-loop breaking, completion veto, reasoning-effort governance and model routing — callable as a CLI, MCP server, git/CI hook and typed SDKs across **three runtimes (Python, TypeScript, Rust)**.
> **What it is not:** a coding agent; a runtime supervisor (that is [Foreman](https://github.com/thruwire/foreman)); a tool-call guardrail ([jev-guard](https://github.com/leepokai/jev-guard)); a context sieve ([Winnow](https://github.com/GhalebDweikat/winnow)); or a capability router ([JevRouter](https://github.com/BillionsBobby/JevRouter)). See [Where it fits](#-where-it-fits-the-system-15-decision-layer).

📑 Table of contents

- [The Problem](#-the-problem) · [How It Works](#-how-it-works-system-1--system-15--system-2) · [Where It Fits](#-where-it-fits-the-system-15-decision-layer) · [Features](#-features) · [Quickstart](#-quickstart) · [CLI](#-cli-usage) · [Astra-Jev effort governance](#-astra-jev-dynamic-reasoning-effort-governance-2026-frontier-models) · [Integrations](#-universal-agent--ide-integrations) · [SDKs](#-python-sdk) ([TS](#-typescript--javascript-sdk--cli), [Rust](#-rust-crate--standalone-cli)) · [Git & CI guardrails](#-git--cicd-guardrails) · [Economics & benchmarks](#-economics--benchmarks-september-2026-frontier) · [Architecture & roadmap](#-architecture--roadmap) · [Contributing](#-contributing--submissions) · [Release](#-multi-registry-release--synchronization-pypi-npm-cratesio) · [Changelog](#-whats-new-in-v020)

---

## ⚡ 15-Second Quickstart

Get running in seconds across any stack — zero external dependencies, zero API keys required for local deterministic mode:

```bash
# 1. Try it instantly without installing (zero token cost, zero API keys needed)
npx @ismaelsoilet/jev-harness triage --sample "ModuleNotFoundError: No module named 'pytest'"

# 2. Or install in your favorite runtime:
pip install jev-harness                 # Python (CLI + SDK)
npm install @ismaelsoilet/jev-harness   # Node.js (CLI + SDK)
cargo install jev-harness               # Rust (Standalone CLI 'jev')

# 3. Verify health and active engine mode:
jev-harness status
```

---

## 🎯 The Problem

When an autonomous coding agent encounters a test failure or compiler error, the standard reaction is to dump 500 lines of raw traceback into an expensive frontier reasoning model (GPT-6 Astra, Claude Fable 5.1). 

| Failure Scenario | Without Jev Harness | With Jev Harness |
| :--- | :--- | :--- |
| **Missing dependency** (`ModuleNotFoundError`, `Cannot find module`, `TS2307`, `E0463`) | 💸 **~50,000 LLM tokens (estimate)** (~$0.50 - $2.50) + 15s delay to output `pip/npm install ...` | ⚡ **Jev triage** (measured:  **📅 Provider verification date: 2026-09-22.** Model IDs, free tiers and prices change weekly — agents and engineers should re-verify them (and record their own date) if more than 30 days have passed. Step-by-step key acquisition for every provider: **[Universal AI Agent Integration Guide](docs/AGENT_INTEGRATION_GUIDE.md#-provider-access--api-keys)**.
>
> **🔒 Privacy:** offline mode (`--mock`, or no credentials) makes **zero network calls**. Live mode transmits the typed questions and the failure log (head 2,000 + tail 4,000 characters) to the provider endpoint. Since v0.2.0 the **state itself is redacted** in all three runtimes: credential-shaped material (API keys, JWTs, GitHub/AWS tokens, DB URLs, private keys) is masked before it leaves the process — in every gate and over MCP/SDK alike. That is shape-based hygiene, **not** a data-loss-prevention layer: structured customer data that does not look like a credential is still transmitted, so use `--mock` for repositories with regulated data.

Credential resolution priority:
1. Environment variables (`TYPESAFE_API_KEY`, `CMD_API_KEY`, `COMMAND_CODE_API_KEY`, `OPENCODE_API_KEY`, `OPENROUTER_API_KEY`, or `AI_GATEWAY_API_KEY`)
2. Local repository `.jev.json`, `.env`, or `~/.commandcode/auth.json`
3. Global configuration `~/.config/jev/credentials.env`
4. **Resilient Fallback**: active when no credentials are configured and on retryable provider failures (`429`/`5xx`/timeouts/network) or `401`/`403` from **any** provider. Retries use capped backoff + `Retry-After`; the engine prints `[JEV WARNING]` to stderr and every degraded result is flagged `is_mock=true` with a `degraded_reason`. Use `--fail-closed` to surface those errors instead (CLI exit `2`, no traceback).

```bash
# Check current connection & provider status anytime
jev-harness status
```

### Repository Configuration (`.jev.json`)

`jev-harness init` scaffolds a repository-local `.jev.json`. Honored keys:

| Key | Type | Default | Effect |
| :--- | :--- | :--- | :--- |
| `model` | string | provider default | Overrides the model sent to the provider. The scaffold placeholder `jev-latest` means "use the provider-optimized default", so it never clobbers provider model IDs. |
| `skip_llm_threshold` | float `0`-`1` | `0.65` | Minimum confidence for `test-gate` to set `skip_llm=true` (a `deep_logic` verdict is never bypassed). |
| `abort_threshold` | float `0`-`1` | `0.70` | Minimum dead-end probability for `abort-check` to abort a trajectory. |
| `shadow` | bool | `false` | Decide and report, but never change the exit code (see [Shadow Mode](#-features)). |
| `api_key` / `provider` | string | — | Optional credentials. Environment variables take precedence. |

Values are clamped to `[0, 1]`, and a corrupted file degrades to defaults instead of breaking CI. The same keys work identically in Python, TypeScript and Rust.

**Pinning the model.** The effective model resolves as explicit argument → `JEV_MODEL` env var → `model` in `.jev.json` → provider default (run `jev-harness status` to see both the model and where it came from). The alias `jev-latest` **moves**: the provider may change what it points to. Once your `skip_llm_threshold` is calibrated against a model version, pin it — for example `"model": "jev-1.13.0"` — so a provider release cannot silently change your decisions.

---

## 🛠️ CLI Usage

### 1. Test Failure Triage (`test-gate`)
Pipe error logs directly or pass a file:

```bash
# Pipe directly from your test runner
npm test | jev-harness test-gate
pytest | jev-harness test-gate

# Or analyze a saved log file
jev-harness test-gate --log error.log

# Or get machine-readable JSON
jev-harness test-gate --log error.log --json

# Observe decisions without blocking a pipeline (always exits 0)
jev-harness test-gate --shadow --log error.log

# Fail hard on provider outages instead of degrading to the offline engine
jev-harness test-gate --fail-closed --log error.log

# Bound how long retryable provider failures are retried (default: 3 attempts)
jev-harness test-gate --retries 1 --log error.log
```

**Output Example:**
```text
--- JEV TEST TRIAGE VERDICT ---
Category:        ENV_MISSING
Confidence:      92.0%
Skip LLM Call:   YES (Save Tokens!)
Skip Probability: 96.0%
Severity Score:  1.0 / 4.0
Recommendation:  AUTO-ACTION: Install missing dependency or check environment configuration (Do NOT call LLM).
--------------------------------
```

### 2. Guard Against Doom Loops & Dead-Ends (`abort-check`)
Verify that a proposed plan isn't repeating a failed path:

```bash
jev-harness abort-check \
  --plan "Retry rewriting the entire database schema without backup" \
  --history "Attempt 1 failed with timeout. Attempt 2 failed with circular foreign key error."
```
*Returns exit code `1` if abort is recommended, enabling automated CI stops.*

### 3. Model Tier Routing (`route`)
Pick the cheapest model capable of solving the task:

```bash
jev-harness route --task "Fix typo in docstring and reformat with black"
# -> TIER: DETERMINISTIC | Model: Direct Python/Bash Script (0 LLM Tokens)

jev-harness route --task "Refactor distributed actor supervision tree across 14 modules"
# -> TIER: HEAVY_SYSTEM2 | Model: Claude Fable 5.1 / GPT-6 Astra (~$10.00 in / $50.00 out)
```

### 4. Step Completion Verification (`verify`)
Verify evidence against criteria with calibrated confidence:

```bash
jev-harness verify \
  --criteria "Must export format_date function and pass all 10 unit tests" \
  --output "All 10 unit tests passed in 0.02s. format_date exported in index.ts."
```

### 5. Dynamic Reasoning Effort Governance (`reasoning-effort` / `astra-jev`)
Dynamically modulate reasoning effort per-generation (inspired by Vechen @miu21590) to eliminate latency and save thousands of tokens on mechanical tool steps:

```bash
# Evaluate immediate step for DeepSeek (e.g. DeepSeek V4.1-Flash / V4-Pro)
jev-harness reasoning-effort \
  --context "git status e verificar arquivos alterados no commit recente" \
  --target-provider deepseek

# Output:
# Effort: LOW | Dialect: {"extra_body": {"thinking": {"type": "enabled"}}, "reasoning_effort": "low"}
# Latency eliminated: ~200s internal CoT reduced to 1.5s!

# Evaluate architectural task for Anthropic (Claude Fable 5.1 / Claude Opus 5)
jev-harness reasoning-effort \
  --context "Architect distributed actor supervision tree with raft consensus" \
  --target-provider anthropic --json

# Safeguard check for direct models (returns empty params and warnings for non-reasoning models)
jev-harness reasoning-effort \
  --context "Run bash command" \
  --target-provider openai \
  --model gpt-5.6-luna
```

### 6. Continuation Nudge Gate (`nudge-gate` / `nudge`)
Inspired by [`CommandCodeAI/cmd-mod-jev-nudge`](https://github.com/CommandCodeAI/cmd-mod-jev-nudge), `nudge-gate` combines gated workflow phases (`research`, `ask`, `plan`, `execute`, `verify`, `complete`) with calibrated `Noul` probabilities (`nudge`, `waiting`, `progress`) to evaluate whether an autonomous agent paused prematurely with unfinished work or unverified changes (`should_nudge = true`, exit code `0`), while automatically vetoing nudges when waiting on user input (`waiting >= 0.5` or `phase == "ask"`), when the previous nudge produced no progress (`progress  should_nudge: true | workflow_phase: "verify" | exit code 0

# Evaluate when waiting on user choice (vetoed automatically)
jev-harness nudge-gate \
  --transcript "Assistant: Which AWS region should I deploy to? Would you like me to proceed?"
# -> should_nudge: false | workflow_phase: "ask" | exit code 1
```

### 7. ROI & Token Savings Telemetry (`metrics`)
Inspect cumulative tokens saved, dollars saved, and doom loops intercepted:

```bash
# View active telemetry
jev-harness metrics

# Reset session telemetry counters
jev-harness metrics --reset
```

**Output Example:**
```text
============================================================
              JEV HARNESS TELEMETRY & ROI
============================================================
Total Triage Interceptions:      14 calls
LLM Frontier Calls Skipped:      11 calls (78.6%)
Abort Guard Stops Triggered:     2 doom loops killed
Deterministic Routes:            6 tasks
Reasoning Effort Modulations:    8 steps (6 low, 2 high)
Estimated Tokens Saved:          422,200 tokens (heuristic estimate)
Estimated Frontier Dollars Saved: $6.12 USD (heuristic estimate)
Assumption Model:                26,200 tokens/$0.31 per intercepted triage; 80,000 tokens/$1.20 per aborted doom loop
============================================================
```

> 📊 **These figures are a planning estimate, not metered usage.** The per-event assumptions are fixed constants (26,200 tokens/$0.31 per intercepted triage, 80,000 tokens/$1.20 per aborted loop). `--json` exposes `estimates_are_heuristic: true` so downstream tooling can label them correctly.

### 8. Self-diagnosis, audit trail and calibration (`doctor` / `receipts` / `replay`)

```bash
# Is my installation healthy? (never prints secrets; --live spends ONE request)
jev-harness doctor
jev-harness doctor --live --json

# What did this repository decide? (append-only, hashes + metadata only)
jev-harness receipts --tail 10
jev-harness receipts --json

# How accurate are the gates? (confusion matrix, P/R/F1, ECE per gate; fails on regression)
jev-harness replay --corpus tests/corpus
```

### 7. One-Command Agent Setup (`init`)
Automatically scaffold MCP configurations for your active agent or IDE:

```bash
# Setup for Cursor
jev-harness init --cursor

# Setup for Antigravity IDE
jev-harness init --antigravity

# Setup git pre-commit hook (detects npm/pytest/cargo; never overwrites an existing hook)
jev-harness init --git

# Setup all supported tools at once
jev-harness init --all
```

---

## ⚡ Astra-Jev: Dynamic Reasoning Effort Governance (2026 Frontier Models)

Inspired by Vechen's ([@miu21590](https://x.com/miu21590)) groundbreaking work on *Astra-Codex* and the **[Astra-Ares](https://github.com/miuuyy/Astra-Ares)** framework, **Astra-Jev** introduces autonomous, per-generation reasoning effort modulation governed by TypeSafe Jev System One.

Instead of locking an entire multi-turn coding session into heavy, slow reasoning (or risking bugs by running exclusively in low reasoning), Astra-Jev evaluates the cognitive demand of the immediate next generation in **(*GPT-6 Astra*, *Claude Fable 5.1*) | **Dollar Cost**($10/1M in, $50/1M out) | Agent burns ~8,000 reasoning tokens ($0.40 - $1.20) just to inspect `git status` or read a file | Injects `effort="low"`, burning only ~300 tokens. **Saves up to $1.15 per mechanical generation.** |
| **Chinese Frontier**(*DeepSeek V4.1-Flash*, *Qwen 3.8 Max*, *Kimi-k3*, *MiMo*) | **Latency & GPU Starvation**(Tokens are cheap, but internal CoT takes 3–5 minutes) | Agent enters 200–300 second internal thinking loop before running a trivial bash command | Disables thinking CoT or sets `effort="low"`. Response delivered in **1.5s instead of 240s**. |

### 🛡️ Critical Safeguards Built into Astra-Jev

1. **Direct Single-Pass Model Safeguard:** Models that do not support internal reasoning (e.g. `gpt-5.6-luna`, `gemini-3.8-live`, `claude-3.5-haiku`) will return fatal **HTTP 400 Bad Request** if reasoning parameters are injected. Astra-Jev automatically detects non-reasoning targets, sets `is_reasoning_supported = False`, and returns clean empty payloads `{}`.
2. **Preservation of `reasoning_content` (DeepSeek multi-turn):** In DeepSeek V4.1-Flash/Pro APIs, stripping `reasoning_content` across multi-turn tool calling can corrupt tool execution. Astra-Jev enforces dialect compliance to preserve thinking structures across turn transitions.
3. **Prompt Cache (KV Cache) Trade-off Advisory:** Toggling reasoning parameters back-and-forth mid-session can invalidate prefix cache on long contexts (>100k tokens). Astra-Jev provides `cache_safe_recommendation` advisories:
   - For pure mechanical actions, use Jev's `skip_llm=true` to execute directly without calling the LLM at all.
   - Keep reasoning effort stable across related sub-steps of a single complex implementation.

---

## 🤖 Universal Agent & IDE Integrations

> 📖 **Looking for a turnkey setup for any project?** Read the [**Universal AI Agent Integration Guide**](docs/AGENT_INTEGRATION_GUIDE.md) (

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [ismaelsoilet](https://github.com/ismaelsoilet)
- **Source:** [ismaelsoilet/jev-harness](https://github.com/ismaelsoilet/jev-harness)
- **License:** MIT
- **Homepage:** https://pypi.org/project/jev-harness/

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** yes
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-ismaelsoilet-jev-harness
- Seller: https://agentstack.voostack.com/s/ismaelsoilet
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
