# Praxia

> Multi-agent orchestrator with cyclic personal-to-org memory · Apache 2.0 · Python 3.11+

- **Type:** MCP server
- **Install:** `agentstack add mcp-praxia-dev-praxia`
- **Verified:** Pending review
- **Seller:** [praxia-dev](https://agentstack.voostack.com/s/praxia-dev)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** Apache-2.0
- **Upstream author:** [praxia-dev](https://github.com/praxia-dev)
- **Source:** https://github.com/praxia-dev/praxia
- **Website:** https://praxia.tools

## Install

```sh
agentstack add mcp-praxia-dev-praxia
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Praxia
[](https://pypi.org/project/praxia/)
[](https://pypi.org/project/praxia/)
[](https://github.com/praxia-dev/praxia/actions/workflows/test.yml)
[](https://x.com/praxia_dev)

🌐 **Live**: [praxia.tools](https://praxia.tools/) (primary, Cloudflare) · [praxia-dev.github.io/praxia](https://praxia-dev.github.io/praxia/) (mirror, GitHub Pages) · [@praxia_dev](https://x.com/praxia_dev) on X

[](https://youtu.be/Z3DFa2saHJg "▶ Click to watch the full 4-min demo on YouTube")

📺 [Watch the 4-minute walkthrough](https://youtu.be/Z3DFa2saHJg) · 🚀 [Quickstart](docs/quickstart.md) · 💬 [Discussions](https://github.com/praxia-dev/praxia/discussions)

### Install Praxia Desktop (Windows 10 / 11 x64)

[](https://apps.microsoft.com/detail/9P9LSR34HZF3)

> **Specialized Multi-Agent Orchestrator with Cyclic Personal/Organizational Memory**
>
> A workflow-specific multi-agent orchestrator that **automatically promotes** individual tacit knowledge into organizational know-how. Built on a 5-layer memory stack with three independent promotion paths.

[](LICENSE)
[](https://www.python.org)
[]()
[]()
[]()
[]()
[]()

> 🔍 Complete feature reference: [docs/FEATURES.md](docs/FEATURES.md)
> 📊 Concrete Before/After tables: [docs/use-cases.md](docs/use-cases.md)

---

## ⬇ **Try Praxia in 30 seconds — no Python, no setup**

### 👉 [**📦 Download Praxia Desktop for Windows (.exe, 174 MB)**](https://github.com/praxia-dev/praxia/releases/latest/download/Praxia.Desktop_0.1.0-52_x64-setup.exe)

Windows 10 / 11 x64 alpha · Tauri + embedded Python sidecar · zero
`pip install`, zero `praxia serve`. Paste an LLM API key (Anthropic /
OpenAI / Azure / Google / Qwen / HF / Ollama) and you're chatting.
Unsigned alpha — SmartScreen will warn on first launch, click
**"More info" → "Run anyway"**.

Other downloads · [`.msi` for managed deployment](https://github.com/praxia-dev/praxia/releases/latest/download/Praxia.Desktop_0.1.0-52_x64_en-US.msi)
· [all releases & notes](https://github.com/praxia-dev/praxia/releases/latest)
· macOS / Linux coming in Phase 1b.

---

## 🎯 Why Praxia?

General-purpose multi-agent frameworks (CrewAI, AutoGen, LangGraph, …) are powerful but stop short on these four problems:

| Problem with existing frameworks | Praxia's approach |
|---|---|
| Setup is complex; production deployment is hard | **Workflow-specific templates** (sales prep / logic check / RAG optimization) that run in 5 minutes |
| Senior-engineer "magic prompts" stay locked in one person's editor | **Personal-to-org auto-promotion pipeline** built in |
| "It works" doesn't prove "it works *well*" | **Hallucination detection + retrieval evals** shipped by default |
| Agents stagnate after launch | **Sleep-time consolidation** distills your past flows nightly |

Praxia turns "one expert's drawer" into "everyone's best practices."

---

## 👥 Who Praxia is for

| Persona | What they need | How Praxia fits | Typical year-1 result |
|---|---|---|---|
| **🏢 Information Systems / Platform team** (300–5,000 employees) | Roll out AI tools without paywalled SSO/RBAC/audit, on-prem option | Auth + RBAC + ACL + per-user OAuth + audit log all in OSS, self-hostable | 100 KW × ~$1.25M net benefit, full audit trail |
| **🏗️ Engineering / Product VP** (50–500 in scope) | Senior architect bottleneck; junior PM 12–18mo ramp | DesignSkill + sleep-time consolidation distills senior review patterns; Markdown+git frozen layer fits PR workflow | Senior load 16h/wk → 4h/wk; junior ramp 6–9mo |
| **⚖️ Legal / Compliance lead** (regulated industry) | 50–100 contracts/mo bottleneck; need *auditable* AI workflow without lock-in | LegalSkill (RACE) + read-only memory mode + per-user OAuth + every action audited; Apache 2.0 source for auditors | 60–90min → 10–15min/contract; throughput 50–80/mo → 200–300/mo |
| **🧪 OSS / Research integrator** | Build domain agent system without re-implementing auth, memory cycling, exporters | 7 plugin types (~50 LoC each); use as library, run `praxia serve` as backend, embed in LangGraph | Day-30: domain skill + custom connector + memory cycling working — **~3 weeks ahead of from-scratch** |

Detailed Before/After by industry: [docs/use-cases.md](docs/use-cases.md).

---

## 💎 Why OSS matters here

The capabilities you typically pay enterprise tier for — already in the Apache 2.0 package:

- **SSO + RBAC + audit are not paywalled.** OIDC SSO (Google / Microsoft / Okta / GitHub / Keycloak) is in the OSS. Most agent frameworks ship without it; most agent platforms paywall it. Praxia treats it as table stakes.
- **Memory format is not locked in.** Layer 4 is plain Markdown in your git repo. Layer 3 exports to JSONL. Layer 1 is your chosen backend's native format. Leaving costs nothing.
- **You can read every line.** Apache 2.0. Show the source to your auditors, your security team, your customers.
- **Multi-LTM ensembles, not single-vendor.** Run Mem0 + Zep + HindSight in parallel, fuse with RRF, or route per query. No commercial agent platform exposes this — they pick a backend and lock you in.
- **Per-user OAuth respects external ACL.** When Alice pulls from Box, Box's own ACL applies — Alice only sees what Alice can see. Service-account designs (typical SaaS shortcut) leak data across users.
- **Air-gapped operation.** `PRAXIA_LOCAL_MODEL=gemma`, Ollama, `backend=json` — no cloud LLM, no cloud vector DB, no telemetry. Same code as cloud customers.
- **Production-grade OAuth + KMS in OSS.** Multi-worker safe state cache, 5 KMS adapters (AWS / Azure / GCP / Vault / local). Most agent platforms paywall this; Praxia ships it.
- **A/B experiments + quality eval included.** Test prompt variants on real users with deterministic assignment; catch LLM output quality regressions in CI.

---

## 🏗 Architecture — 5-Layer Memory Stack

The same picture as ASCII art:

```
┌──────────────────────────────────────────────────────────┐
│  AI Agents (Skills + MCP)                                │
└──────────────┬───────────────────────────────────────────┘
               │ Users just have normal conversations
               ▼
╔═══════════════════════════════════════════════════════════╗
║ Layer 1: Personal memory (auto-extracted)                 ║
║   Mem0 / LangMem / HindSight / Letta / Zep / JSON         ║
║   namespace = user_id                                     ║
║   ★ Zero-effort tacit-knowledge capture                   ║
╚══════════════╤════════════════════════════════════════════╝
               │ Sleep-time Consolidation (nightly batch)
               ▼
╔═══════════════════════════════════════════════════════════╗
║ Layer 2: Distillation & promotion engine                  ║
║   Three parallel "validity tests":                        ║
║     ① Frequency  (recurring across N+ users)              ║
║     ② Outcome    (correlated with wins/losses)            ║
║     ③ Self-eval  (LLM scored)                             ║
╚══════════════╤════════════════════════════════════════════╝
               │ Auto-promote above threshold; queue otherwise
               ▼
╔═══════════════════════════════════════════════════════════╗
║ Layer 3: Shared memory (living organizational knowledge)  ║
║   Letta-style shared blocks; all agents read/write        ║
╚══════════════╤════════════════════════════════════════════╝
               │ PR review for high-impact items
               ▼
╔═══════════════════════════════════════════════════════════╗
║ Layer 4: Frozen layer (git-managed best practices)        ║
║   Markdown + git + PR review                              ║
║   GitHub Copilot / Cursor Rules-compatible format         ║
╚══════════════╤════════════════════════════════════════════╝
               │ (optional)
               ▼
╔═══════════════════════════════════════════════════════════╗
║ Layer 5: Graph layer (only relationship-heavy domains)    ║
║   Zep / Graphiti — decisions, customer 360, incident DAG  ║
╚═══════════════════════════════════════════════════════════╝

Parallel Layer 6: Skills registry
  Personal skills get promoted to the organizational catalog.
  MCP / Claude Skills / Cursor Skills compatible.
```

Three promotion paths (**auto / statistical / manual**) run side by side — never depending on a single mechanism.

For details, see [docs/architecture.md](docs/architecture.md).

---

## ✨ What's Bundled

### Autonomous agent (LLM-driven tool-use loop)

`praxia.agent.AutonomousAgent` runs an LLM-driven tool-use loop over the
full Praxia stack — personal/org memory, skills, frozen layer, connectors —
with ACL checks and audit logging built in. The LLM picks tools on its own
until it has the information it needs, mirroring how modern code-editing assistants drives its
own tool use.

```python
from praxia.agent import AutonomousAgent
from praxia.core.llm import LLM

agent = AutonomousAgent(user_id="alice", org_id="acme", llm=LLM("claude"))
result = agent.run("Tell me what we know about Acme and draft a proposal.")
print(result.final_text)
```

```bash
praxia agent run "Summarize where we stand with Acme this quarter and draft a proposal"
praxia agent tools     # list the 15 built-in tools
```

The agent is also exposed as a single MCP meta-tool (`autonomous_agent`) so
remote clients (Claude Desktop, Cursor) can delegate an entire investigation
without orchestrating individual tools by hand. See
[FEATURES § 38](docs/FEATURES.md#38-autonomous-agent-llm-driven-tool-use-loop).

### CommandedAgent — autonomous agent with external verification

`AutonomousAgent` is a free-running tool-use loop — perfect when the
environment *is* the answer key (tests pass / fail, commands exit 0 /
non-zero). For workloads where the environment doesn't give you that
free check — private-corpus fact QA, SOP / compliance, customer
support over manuals, technical-knowledge transfer — `CommandedAgent`
wraps it with three guards: **pre-retrieval + grounding verification +
bounded retry**, with an explicit `abstain` path when the sources
don't support a confident answer.

Calibrated against an in-house multi-hop RAG harness
(HotpotQA / SQuAD v2 / JEMHopQA) — see
[`docs/VERIFICATION_FINDINGS.md`](docs/VERIFICATION_FINDINGS.md) — and
the resulting defaults:

- **Task-type router** — `default_task_classifier` sends coding /
  command / tool prompts straight through (the environment is the
  verifier for those), and routes knowledge-QA through the grounding
  gate. Bilingual (EN/JA) keyword classifier; pluggable.
- **Query decomposition** — wire a `QueryDecomposer` into the
  retriever and multi-hop questions are split into sub-questions,
  retrieved per hop, and unioned (deduped). +12pt on HotpotQA-distractor
  40q in the source benchmark; single-hop queries pass through
  unchanged.
- **No-improvement early stop** — a redraft that fails to lift
  groundedness by `min_groundedness_improvement` (default 0.05) aborts
  to abstain instead of burning the whole round budget on a loop that
  isn't making progress.
- **Promotion engine hardened** — `PromotionEngine` blends
  `0.5·freq + 0.4·outcome + 0.1·self_eval` and **alpha26+ adds two
  hard pre-condition gates that short-circuit the score blend**:
  (1) an optional external `verifier(text, verdict) → (passed,
  reason)` callback — `passed=False` forces `decision="skip"` no
  matter how high the blend scored, closing the self-eval reward-
  hacking path identified in the L0–L3 verification bench; (2)
  `min_independent_sources=2` — single-source candidates cap at
  `"review"` and never auto-promote, blocking the
  contamination-from-one-user path. `require_verification=True`
  fails closed when no verifier is wired in.

```python
from praxia.agent import AutonomousAgent, CommandedAgent
from praxia.agent.decomposer import LLMQueryDecomposer
from praxia.agent.commander import DefaultMemoryRetriever
from praxia.core.llm import LLM

llm = LLM("claude")
inner = AutonomousAgent(user_id="alice", org_id="acme", llm=llm)
retriever = DefaultMemoryRetriever(
    personal=inner.memory,                  # L1
    decomposer=LLMQueryDecomposer(llm=llm), # K2: multi-hop split
)
agent = CommandedAgent(inner, retriever=retriever, max_verify_rounds=3)

result = agent.run(
    "How do we handle Customer X's stamping-press alarm code E-204?"
)
# result.answer carries [L1#0, L3#2, ...] citations
# result.verdict.decision ∈ {"accept", "redraft", "abstain"}
# result.stopped_reason ∈ {"accept", "abstain", "no_improvement",
#                          "max_rounds", "bypass_action"}
# result.task_kind         ∈ {"knowledge", "action"}
# result.rounds            is the per-round draft + verdict trace
```

Pluggable everything: `Verifier` / `Retriever` / `QueryDecomposer` /
`TaskClassifier` are all protocols or callables, so you can drop in
TiDB Vector / pgvector / hybrid BM25+ANN / GraphRAG or your own
grounding scorer / decomposer without touching core. See
[FEATURES § 38b](docs/FEATURES.md#38b-commandedagent--autonomous-agent-with-external-grounding-commander)
and [`docs/COMMANDED_AGENT.md`](docs/COMMANDED_AGENT.md).

### 3 Specialized Multi-Agent Flows

| Flow | What it does |
|---|---|
| **SalesAgentFlow** | Reads customer IR, past minutes, RAG context → generates **hypotheses → FAQ → proposal outline** |
| **LogicCheckerFlow** | Three agents (structure / contradiction / reader) review long documents for logical consistency |
| **RAGOptimizationFlow** | Self-correcting RAG: query expansion → retrieval → relevance eval → hallucination check loop |

### 6 Default Business-Domain Skills

| Skill | Domain | Use cases |
|---|---|---|
| **InvestmentSkill** | Investment | Equity research, due diligence, portfolio decisions |
| **SalesSkill** | Sales | Account research, proposal drafting, FAQ prep |
| **DesignSkill** | Engineering Design | System design review, requirements engineering |
| **PurchasingSkill** | Procurement | Supplier evaluation, RFQ analysis, TCO, BCP risk |
| **PatentSkill** | IP / Patent | Prior-art search, claims drafting, patent maps |
| **LegalSkill** | Legal | Contract review, compliance, M&A diligence |

Each skill serializes to Claude-Skills / MCP-compatible `SKILL.md`.

Plus four **utility** skills:

| Skill | What it does |
|---|---|
| **`PromptDesignerSkill`** | Take a one-line task description → produce a production-grade prompt template (system + user + 2-3 few-shot examples + 5-criterion rubric) tuned for the target LLM (Claude / OpenAI / DeepSeek / Mistral / Llama / …). Save to `PromptStore`, A/B-test via `praxia.experiments`. |
| **`OutputFormatSkill`** | Detect "export as PowerPoint" / "as Word doc" / etc. in natural language (English + JA / ZH / KO / ES / FR / DE / PT-BR phrases also recognized) and dispatch to the matching exporter (PPTX / DOCX / HTML / MD / JSON). |
| **`PptxDesignerSkill`** | **Code-gen** path (Claude-Skills-style): the LLM authors `python-pptx` code that runs in a sandbox to produce a design-rich `.pptx` — multi-column layouts, matrix slides, embedded charts, themed branding. Themes (colors / fonts / logo / footer) live under `.praxia/themes//` and are managed in `Admin → 🎨 Themes`. |
| **`DocxDesignerSkill`** | Same approach for Word documents — LLM-authored `python-docx` code, sandbox execution, themed `.docx` output (heading hierarchy, page footer, tables, callouts, embedded charts). |

```bash
# Generate a prompt template for any task
praxia skill run prompt_designer "Have in-house legal score contract risk on a 5-point scale"

# Generate a design-rich slide deck via code-gen
praxia skill run pptx_designer "Q4 sales review for Acme — cover, exec summary, top 3 customers, 2x2 challenges matrix, next actions. 10 slides."
```

### All Major LLMs

LiteLLM-powered single-line provider switching:

| Provider | Aliases (current default) | Auth env var(s) |
|---|---|---|
| Anthropic Claude | `claude` (Opus 4.7) · `claude-sonnet` (4.6) · `claude-haiku` (4.5) | `ANTHROPIC_API_KEY` |
| OpenAI ChatGPT | `chatgpt` (GPT-5.5) · `gpt-5.5` · `gpt-5.4` · `gpt-5` · `o4-mini` · `o3` · `gpt-4o` | `OPENAI_API_KEY` |

…

## Source & license

This open-source MCP server is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [praxia-dev](https://github.com/praxia-dev)
- **Source:** [praxia-dev/praxia](https://github.com/praxia-dev/praxia)
- **License:** Apache-2.0
- **Homepage:** https://praxia.tools

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** yes
- **Dynamic code execution:** yes

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: flagged — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/mcp-praxia-dev-praxia
- Seller: https://agentstack.voostack.com/s/praxia-dev
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
