# Pydantic Deepagents

> Open-source, self-hosted Claude Code - a terminal AI assistant and the Python framework behind it. Tool-calling, sandboxed execution, multi-agent teams, skills, checkpoints, unlimited context - on Pydantic AI, any model.

- **Type:** MCP server
- **Install:** `agentstack add mcp-vstorm-co-pydantic-deepagents`
- **Verified:** Pending review
- **Seller:** [vstorm-co](https://agentstack.voostack.com/s/vstorm-co)
- **Installs:** 0
- **Category:** [Cloud & Infrastructure](https://agentstack.voostack.com/c/cloud-infrastructure)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [vstorm-co](https://github.com/vstorm-co)
- **Source:** https://github.com/vstorm-co/pydantic-deepagents
- **Website:** https://vstorm-co.github.io/pydantic-deepagents/

## Install

```sh
agentstack add mcp-vstorm-co-pydantic-deepagents
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

Pydantic Deep Agents

  Open-source Claude Code — that you can also build on.
  A self-hosted terminal AI assistant and the Python framework behind it.
  Use it today, or ship your own agent in one function call. Any model. 100% type-safe. MIT.

  

  Docs &middot;
  PyPI &middot;
  Forking &middot;
  Why &middot;
  CLI &middot;
  Framework &middot;
  Examples

  
  
  
  
  
  
  
  
  
  
  

---

**Pydantic Deep Agents is two things in one repo:**

🖥️ **A terminal AI assistant** — a self-hosted, open-source alternative to Claude Code. Install it, point it at any model, and it plans, edits files, runs commands, searches the web, remembers across sessions, spawns sub-agents, and connects to MCP servers. Almost everything Claude Code does — on the model *you* choose.

🐍 **A Python framework** — the *exact same harness* behind a single function call. `create_deep_agent()` hands a model a filesystem, shell, planning, memory, sub-agents, sandboxed execution, MCP, and unlimited context. Build your own assistant, research agent, or coding tool without rewiring the plumbing every time.

Both run on **[Pydantic AI](https://github.com/pydantic/pydantic-ai)**, work with **any model** (Claude, GPT, Gemini, local), and are **100% type-safe** and MIT-licensed — and they share one trick nothing else has: [**Live Run Forking**](#-live-run-forking--the-feature-no-one-else-has), splitting a single run into parallel branches an AI judge merges back together.

## Two ways to use it

### 🖥️ 1. Use the assistant

A Claude-Code-style TUI in your terminal, on **any** model — no Python setup (the script installs `uv` + the CLI for you):

```bash
curl -fsSL https://raw.githubusercontent.com/vstorm-co/pydantic-deep/main/install.sh | bash
pydantic-deep
```

> Windows / manual: `pip install "pydantic-deep[cli]"`

### 🐍 2. Build your own

One function call gives you a full deep agent:

```bash
pip install pydantic-deep
```

```python
from pydantic_deep import create_deep_agent

agent = create_deep_agent(model="anthropic:claude-sonnet-4-6")
result = await agent.run("Build a REST API for auth")
```

---

## ⑂ Live Run Forking — the feature no one else has

Claude Code can't do this. Aider can't. LangGraph and CrewAI can't. **It's the reason to use pydantic-deep.**

When an agent hits a fork in the road — "should I refactor this with a decorator or a context manager?" — most tools force one bet. Pydantic Deep Agents lets the run **branch**:

```
                                  ┌──  branch A: "use a decorator"      ── tests: 8/8 ✓  conf 0.71
   agent.run("refactor auth") ──┬─┼──  branch B: "use a context manager" ── tests: 6/8 ✗  conf 0.42
       (shared history)         │ └──  branch C: "extract a base class"   ── tests: 8/8 ✓  conf 0.55
                                │
                                └──►  ⚖️  AI judge weighs quality + tests + consistency
                                          → adopts branch A, continues the run
```

Each branch is **fully isolated**: a copy-on-write filesystem overlay (reads fall through to the parent, writes stay local), its own steering message, and its own `budget_usd` cap. The coordinator resolves the fork with one of four acceptance modes — `manual`, `auto`, `auto_with_fallback` (default), or `vote` — and the winning branch's history is adopted as the parent run's continuation.

**Framework — opt in with one flag:**

```python
agent = create_deep_agent(
    model="anthropic:claude-sonnet-4-6",
    forking=True,                 # gives the agent: fork_run, inspect_branches,
)                                 # merge_or_select, diff_branches, fork_cost, terminate_branch
```

**Or run a real test command against every branch and let exit codes decide the winner:**

```python
from pydantic_deep import LiveForkCapability

agent = create_deep_agent(
    forking=LiveForkCapability(test_command="pytest -q", test_timeout_s=120),
)
# confidence = quality_spread·0.4 + test_pass_ratio·0.4 + internal_consistency·0.2
```

**CLI — fork an in-flight conversation, watch branches stream live, merge the best:**

```
/fork                 # split the current run into N parallel branches
>>A try a decorator   # steer branch A
>>B use a contextmgr  # steer branch B
/merge                # resolve — manual picker, AI judge, or vote
```

Live per-branch panels stream each approach side by side; a judge screen scores them; you accept, review the diff, or decline. Configure branch count, budgets, per-branch models, and merge strategy with `/fork-config`.

> 📖 Full reference: [docs/capabilities/live-fork.md](docs/capabilities/live-fork.md)

---

## 🆚 Why pydantic-deep?

The only tool that is **a terminal assistant** *and* **a Python framework** *and* can **fork its own runs** — without giving up type safety or your choice of model.

| | **Pydantic&nbsp;Deep** | Claude&nbsp;Code | Aider | LangGraph | CrewAI |
|---|:---:|:---:|:---:|:---:|:---:|
| Terminal TUI assistant | ✅ | ✅ | ✅ | — | — |
| Python framework / library | ✅ | — | ~ | ✅ | ✅ |
| **Live run forking + AI judge** | ✅ | — | — | — | — |
| Multi-agent swarm + message bus | ✅ | ~ | — | ✅ | ✅ |
| Any model / any provider | ✅ | Anthropic | ✅ | ✅ | ✅ |
| Sandboxed Docker execution | ✅ | — | ~ | DIY | DIY |
| Persistent memory + skills | ✅ | ✅ | — | DIY | ~ |
| Type-safe structured output | ✅ | — | — | ~ | ~ |
| MCP servers | ✅ | ✅ | — | ~ | ~ |
| Self-hosted, open source | ✅ MIT | — | ✅ | ✅ | ✅ |

✅ first-class · ~ partial / via extensions · — not available · DIY you wire it yourself. Comparison reflects each project as of 2026-06; corrections welcome via PR.

---

## What's New

- **2026-06-01** &nbsp;**v0.3.24** — **Live Run Forking** — split an in-flight `agent.run()` into N parallel branches with copy-on-write isolation, per-branch budgets, a test-runner hook, and four merge modes (`manual` / `auto` / `auto_with_fallback` / `vote`). Opt in with `forking=True`.
- **2026-06-01** &nbsp;**v0.3.23** — **MCP client support** (framework + CLI). Connect GitHub, Figma (OAuth), Context7, DeepWiki, or any custom server. Import servers straight from Claude Code. New interactive `/mcp` command. Plus a full CLI presentation pass: clipboard image paste, real `+/-` diffs, tool icons, turn summaries.
- **2026-06-01** &nbsp;**v0.3.23** — **Automatic fallback-model retry** — `fallback_model=` wraps your primary in a `FallbackModel` chain; fires on API errors but never on auth errors. Plus a batteries-included **security hook preset** (`default_security_hook()`) and three new output styles (`markdown`, `json-only`, `bullet`).
- **2026-04-22** &nbsp;**v0.3.17** — LiteParse document parsing (`include_liteparse=True`) — PDFs, DOCX, XLSX, PPTX, and images with optional OCR, all local.
- **2026-04-10** &nbsp;**v0.3.5** — Headless runner (`pydantic-deep run`), Docker sandbox with named workspaces, browser automation via Playwright.

> Full history: [CHANGELOG.md](CHANGELOG.md)

---

## The Agent Harness

Pydantic Deep Agents is an **agent harness** — the complete infrastructure that wraps an LLM and makes it a functional autonomous agent. The model provides intelligence; the harness provides planning, tools, memory, sandboxed execution, unlimited context, and — uniquely — the ability to fork.

⑂ Live run forking
Split a run into N isolated branches, each trying a different approach. AI judge or test results pick the winner. No other agent framework has this.

🔧 Tool-calling
File read/write/edit, shell execution, glob, grep, web search, web fetch, browser automation — wired up and ready.

🤝 Multi-agent / swarm
Spawn subagents for parallel workstreams. Shared TODO lists with claiming. Peer-to-peer message bus. Full team coordination.

🧠 Persistent memory
MEMORY.md persists across sessions. Auto-injected into the system prompt. Each agent has isolated memory by default.

♾️ Unlimited context
Auto-summarization when approaching the token budget. LLM-based or zero-cost sliding window. Never hits a context wall.

🐳 Sandboxed execution
Docker sandbox with named workspaces. Installed packages persist between sessions. Project dir mounted at /workspace.

🗂️ Plan Mode
Dedicated planner subagent asks clarifying questions and structures the work before execution begins. Headless-compatible.

🔖 Checkpoints
Save conversation state at any point. Rewind to any checkpoint. Fork sessions to explore alternative approaches.

📚 Skills system
Domain-specific knowledge loaded on demand from SKILL.md files. Built-in: code-review, refactor, test-writer, git-workflow, and more.

📄 Document parsing
Parse PDFs, DOCX, XLSX, PPTX, and images with optional OCR via LiteParse. Runs locally — no cloud services required.

🔌 MCP
Connect any Model Context Protocol server — GitHub, Figma (OAuth), Context7, DeepWiki, or custom. Import straight from Claude Code.

⚡ Lifecycle hooks + security preset
Claude Code-style PRE/POST_TOOL_USE hooks. Shell or Python handlers. default_security_hook() blocks destructive commands out of the box.

📐 Structured output
Type-safe Pydantic model responses via output_type. No JSON parsing. No dict["key"]. Full IDE autocomplete.

🔁 Fallback models
Primary model fails? fallback_model= hops to the next in the chain — on API errors, never on auth errors.

🔄 Stuck loop detection
Detects repeated identical tool calls, A-B-A-B alternating patterns, and no-op calls. Warns the model or stops the run.

💰 Cost tracking
Real-time token and USD cost tracking per run and cumulative. Hard budget limits with BudgetExceededError.

✨ Self-improving
/improve analyzes past sessions and proposes updates to MEMORY.md, SOUL.md, and AGENTS.md.

🏷️ 100% type-safe
Pyright strict + MyPy strict. 100% test coverage. Every public API is fully typed — safe to use in production.

Built natively on **[pydantic-ai](https://github.com/pydantic/pydantic-ai)** — uses the Capabilities API directly, inherits all pydantic-ai streaming, multi-model support, and Pydantic validation automatically.

---

## 🖥️ CLI — Terminal AI Assistant

A Claude Code-style terminal AI assistant that works with **any model and any provider** — and forks.

### Install (macOS & Linux)

```bash
curl -fsSL https://raw.githubusercontent.com/vstorm-co/pydantic-deep/main/install.sh | bash
```

No Python setup required — the script installs uv and the CLI automatically. Then:

```bash
export ANTHROPIC_API_KEY=sk-ant-...
pydantic-deep
```

> **Windows / manual:** `pip install "pydantic-deep[cli]"` &nbsp;·&nbsp; **Update:** `pydantic-deep update`

### Model & Provider Support

Works with any model that supports tool-calling:

| Provider | Example models |
|----------|----------------|
| **Anthropic** | `anthropic:claude-opus-4-6`, `claude-sonnet-4-6` |
| **OpenAI** | `openai:gpt-5.4`, `gpt-4.1` |
| **OpenRouter** | `openrouter:anthropic/claude-opus-4-6` (200+ models) |
| **Google Gemini** | `google-gla:gemini-2.5-pro` |
| **Ollama (local)** | `ollama:qwen3`, `ollama:llama3.3` |
| **Any OpenAI-compatible** | Custom base URL via env |

Switch model anytime: `pydantic-deep config set model openai:gpt-5.4` or `/model` in the TUI.

### What you get in the TUI

| | Feature |
|:-:|---------|
| ⑂ | **Live run forking** — split a run into branches, stream them side by side, merge the winner |
| 💬 | Streaming chat with tool call visualization, icons, and real `+/-` diffs |
| 📁 | File read / write / edit, shell execution, glob, grep |
| 🤝 | Task planning, plan mode, and subagent delegation |
| 🧠 | Persistent memory and self-improvement across sessions |
| ♾️ | Context compression for unlimited conversations |
| 🔖 | Checkpoints — save, rewind, and fork any session |
| 🔌 | MCP servers via `/mcp` — GitHub, Figma (OAuth), and more; import from Claude Code |
| 🌐 | Web search & fetch built-in · 🖥️ browser automation via Playwright (`--browser`) |
| 🐳 | Docker sandbox — sandboxed execution with named workspaces |
| 💭 | Extended thinking — `minimal` / `low` / `medium` / `high` / `xhigh` |
| 📋 | Clipboard image paste (`Ctrl+V` / `/paste`) — multimodal prompts |
| 💰 | Real-time cost and token tracking per session |
| 🛡️ | Tool approval dialogs — approve, auto-approve, or deny per tool call |
| @ | `@filename` file references · `!command` shell passthrough |
| ✨ | `/fork`, `/merge`, `/improve`, `/skills`, `/mcp`, `/model`, `/theme`, `/compact`, and more |

### Usage

```bash
# Interactive TUI (default)
pydantic-deep
pydantic-deep tui --model openrouter:anthropic/claude-opus-4-6

# Headless deep agent — benchmarks, CI/CD, scripted automation
pydantic-deep run "Fix the failing test in test_auth.py"
pydantic-deep run --task-file task.md --json

# Docker sandbox — sandboxed execution, project dir mounted at /workspace
pydantic-deep tui --sandbox docker
pydantic-deep tui --workspace ml-env     # named workspace, packages persist

# Browser automation (requires pydantic-deep[browser])
pydantic-deep tui --browser

# Config & skills
pydantic-deep config set model anthropic:claude-sonnet-4-6
pydantic-deep skills list
pydantic-deep update                     # update to latest version
```

> See [CLI docs](docs/cli/index.md) for the full reference.

---

## 🐍 Framework — Build Your Own Agent

```bash
pip install pydantic-deep
```

One function call gives you a production deep agent with planning, tool-calling, multi-agent delegation, persistent memory, unlimited context, forking, and cost tracking. Everything is a toggle:

```python
from pydantic_ai_backends import StateBackend
from pydantic_deep import create_deep_agent, create_default_deps

agent = create_deep_agent(
    model="anthropic:claude-sonnet-4-6",
    forking=True,               # ⑂ split a run into parallel branches + AI judge
    include_todo=True,          # Task planning with subtasks and dependencies
    include_subagents=True,     # Multi-agent swarm — delegate to subagents
    include_skills=True,        # Domain-specific skills from SKILL.md files
    include_memory=True,        # Persistent memory across sessions
    include_plan=True,          # Structured planning before execution
    include_teams=True,         # Agent teams with shared TODO lists + message bus
    include_liteparse=True,     # Document parsing — PDF, DOCX, XLSX + OCR
    web_search=True,            # Tool-calling: web search
    thinking="high",            # Extended thinking / reasoning effort
    context_manager=True,       # Unlimited context via auto-summarization
    cost_tracking=True,         # Token/USD budget enforcement
    fallback_model="openai:gpt-5.4",   # auto-retry if the primary model fails
    include_checkpoints=True,   # Save, rewind, and fork conversations
)

deps = create_default_deps(StateBackend())
result = await agent.run("Build a REST API for user auth", deps=deps)
```

### Structured Output

Type-safe responses with Pydantic models — no JSON parsing, no `dict["key"]`:

```python
from pydantic import BaseModel

class CodeReview(BaseModel):
    summary: str
    issues: list[str]
    score: int

agent = create_deep_agent(output_type=CodeReview)
result = await agent.run("Review the auth module", deps=deps)
print(result.output.score)  # fully typed
```

### Multi-Agent Swarm

Spawn isolated subagents for parallel workstreams. Each subagent is a full deep agent with its own tool-calling, memory, and context:

```python
agent = create_deep_agent(
    subagents=[
        {
            "name": "researcher",
            "description": "Researches topics using web search",
            "instructions": "Search the web, synthesize findings, cite sources.",
        },
        {
            "name": "code-reviewer",
            "description": "Reviews code for quality, security, and performance",
            "instructions": "Check for security issues, N+1 queries, missing tests...",
        },
    ],
)
# Main agent delegates: task(description="Review auth.py", subagent_type="code-reviewer")
```

### Claude Code-Style Lifecycle Hooks + Security Preset

```python
from pydantic_deep import create

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [vstorm-co](https://github.com/vstorm-co)
- **Source:** [vstorm-co/pydantic-deepagents](https://github.com/vstorm-co/pydantic-deepagents)
- **License:** MIT
- **Homepage:** https://vstorm-co.github.io/pydantic-deepagents/

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** yes
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: flagged — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-vstorm-co-pydantic-deepagents
- Seller: https://agentstack.voostack.com/s/vstorm-co
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
