AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Agent Architecture Planner

skill-othmane-khadri-gtm-engineer-playbook-agent-architecture-planner · by Othmane-Khadri

Use when designing an autonomous agent, planning agent architecture, building a scheduled automation, or creating a Claude Code agent workflow. Triggers: 'design an agent', 'build an automation', 'agent architecture', 'automate this workflow', 'create a scheduled agent', 'shell script agent'.

No reviews yet
0 installs
26 views
0.0% view→install

Install

$ agentstack add skill-othmane-khadri-gtm-engineer-playbook-agent-architecture-planner

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access Used
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-othmane-khadri-gtm-engineer-playbook-agent-architecture-planner)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
4mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Agent Architecture Planner? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Agent Architecture Planner

Design autonomous agents using the "Shell Orchestrates, Claude Thinks" pattern. This skill takes a workflow description, selects the right architecture pattern, and produces a complete agent package: shell wrapper, system prompt, config, scheduling setup, and cost estimate.

The core principle: the shell handles everything deterministic (API calls, file I/O, scheduling, budget tracking), and Claude handles everything that requires judgment (scoring, filtering, writing, categorizing). This separation makes agents cheaper, more reliable, and easier to debug.


Tools Used

| Tool | Purpose | |---|---| | Read | Check for existing agent docs, load reference files | | Write | Output all agent architecture files | | Bash | Make shell scripts executable, verify directory structure | | Glob | Check for existing agents to avoid naming conflicts |


Methodology

Follow these steps in order. Do not skip steps. Do not produce output files before completing the discovery interview and confirming the architecture with the user.


Step 1 — Workflow Discovery (Interactive)

Present these questions to the user all at once in a numbered list. Tell the user they can answer inline or paste a block of text.

Ask exactly these questions:

I need to understand the workflow before designing the agent. Answer these 6 questions — be as specific as possible.

1. What task do you want to automate? (Describe the manual workflow as you do it today, step by step.)
2. How often should it run? (Every hour, daily, every 2 days, weekly, on-demand only?)
3. What data sources does it need? (APIs, websites to scrape, databases, local files, RSS feeds?)
4. What should it produce? (Report files, database entries, API calls, Slack messages, emails, CSV exports?)
5. What decisions does it need to make? (Scoring relevance, filtering noise, writing content, categorizing items, choosing actions?)
6. What's your budget constraint? (Max cost per run in dollars, or monthly ceiling. Include both API/data costs and LLM costs if you know them.)

Rules for this step:

  • Wait for the user to respond before proceeding. Do NOT generate placeholder answers.
  • If question 1 is vague (less than 3 concrete steps described), ask a focused follow-up: "Walk me through the last time you did this manually. What did you open first? What did you check? What did you produce at the end?"
  • If question 3 mentions an API, ask: "Does this API require authentication? What format does it return (JSON, XML, CSV)? Is there a rate limit?"
  • At most 2 follow-up rounds before proceeding.

Step 2 — Architecture Pattern Selection

Based on the user's answers, recommend one of these four patterns. Present the recommended pattern and briefly explain why it fits. If two patterns are close, present both and let the user choose.


Pattern A: Simple Pipeline
Shell (acquire data) → JSON temp file → Claude (analyze + act) → Output

Best for: Single data source, single output. The most common pattern. Use this when the workflow is: get data, think about it, produce something.

Examples: Daily report generation, email draft writing from CRM data, log analysis, document summarization.

When to use:

  • One data source (or multiple sources that can be merged trivially)
  • No feedback loop needed
  • Output is a single artifact (file, database push, message)

Pattern B: Dual-Layer Pipeline
Shell → Layer 1 (broad collection) → Layer 2 (targeted search) → Merge + Dedup → Claude (analyze + act) → Output

Best for: Monitoring and discovery workflows where comprehensive coverage matters. Layer 1 casts a wide net (e.g., scrape all new posts from 20 sources). Layer 2 does targeted searches for specific keywords or patterns. Results are merged and deduplicated before Claude sees them.

Examples: Social media monitoring, job posting tracking, competitor mention detection, market signal scanning.

When to use:

  • You need to catch everything (missing a relevant item is costly)
  • Data comes from overlapping sources that produce duplicates
  • The volume of raw data is high but only 5-20% is relevant

Pattern C: Multi-Agent Decomposition
Orchestrator Shell → [Agent 1: Research] → [Agent 2: Analysis] → [Agent 3: Writing] → Merge → Output

Best for: Complex workflows where different steps need fundamentally different expertise or prompting strategies. Each sub-agent has its own system prompt optimized for its specific task.

Examples: Content production pipelines (research → outline → draft → edit), multi-step analysis (collect data → score → generate recommendations → format report), workflows where one step's output determines the next step's approach.

When to use:

  • A single system prompt cannot cover all the judgment needed
  • The workflow has 3+ distinct phases that require different thinking
  • You need to be able to debug or improve each phase independently

Pattern D: Feedback Loop
Shell → Collect → Claude (score + filter) → Act → Log results → [wait] → Shell reads logs → Collect → Claude (score with history) → ...

Best for: Ongoing optimization workflows where each run should learn from previous runs. The shell maintains a log of past decisions and outcomes, and feeds them back to Claude as context on subsequent runs.

Examples: A/B testing content strategies, lead scoring with win/loss feedback, adaptive outreach sequences, quality improvement loops.

When to use:

  • Performance should improve over time
  • You have a feedback signal (open rates, conversion, engagement, human ratings)
  • The agent runs repeatedly on similar data

Present your recommendation to the user:

"Based on your workflow, I recommend Pattern {X}: {Name}. Here's why: {1-2 sentence justification}. Does this feel right, or should we look at an alternative?"

Wait for confirmation before proceeding.


Step 3 — Component Design

For the confirmed pattern, design each of the four components below. Present the designs to the user as you go. Do not write files yet — this step is about design, not output.


3A: Shell Wrapper Design

The shell wrapper is the backbone. It handles everything deterministic. Design it with these sections:

Environment Setup

  • Source API keys from environment variables (never hardcode secrets)
  • Define paths (log directory, temp directory, prompt file, config file)
  • Create directories if they don't exist
  • Set timestamp for the current run

Gate Check (prevents double-runs)

GATE_FILE="logs/${AGENT_NAME}_last_run.txt"
MIN_HOURS_BETWEEN_RUNS=44  # adjust per agent

if [ -f "$GATE_FILE" ]; then
    last_run=$(cat "$GATE_FILE")
    now=$(date +%s)
    hours_since=$(( (now - last_run) / 3600 ))
    if [ "$hours_since" -lt "$MIN_HOURS_BETWEEN_RUNS" ]; then
        echo "Gate check: only ${hours_since}h since last run (min: ${MIN_HOURS_BETWEEN_RUNS}h). Skipping."
        exit 0
    fi
fi

Data Acquisition

  • API calls using curl with error handling
  • Write raw responses to temp files (NEVER store API responses in shell variables — they often contain unescaped control characters that break the shell)
  • Validate that the response is not empty or an error

Data Preprocessing

  • Strip JSON to essential fields before sending to Claude. This is the single biggest cost optimization. Example: a raw API response might be 12KB per item. After stripping to the 8-10 fields Claude actually needs, it drops to ~750 bytes per item — a 94% reduction in input tokens.
  • Use a small Python or jq script to do the stripping. Write the stripped data to a new temp file.

Claude Invocation

  • Pipe the prompt file to Claude: cat "$PROMPT_FILE" | claude -p --model {model} --max-turns {turns} --allowedTools {tools}
  • NEVER use claude -p "$(cat ...)" — this hits shell ARG_MAX limits with large prompts (150+ items)
  • Pass the scraped data path as a section in the prompt file, or prepend it to the prompt

Budget Tracking

SPEND_LOG="logs/apify_spend_${AGENT_NAME}.json"
# After each API call, log the cost
# Between batches, check cumulative spend against the daily cap

Logging and Cleanup

  • Log start time, end time, items processed, items passed to Claude, and estimated cost
  • Update the gate file with the current timestamp
  • Clean up temp files

Error Handling

  • If the API call fails (non-200 status), log the error and exit with a non-zero code
  • If Claude returns an error, log it and exit
  • If the data is empty (nothing to process), log "no data" and exit cleanly (code 0)

Present the shell wrapper design as a bulleted outline. Ask the user if any sections need adjustment.


3B: System Prompt Design

The system prompt follows a strict structure. Design it section by section:

1. Role Definition One paragraph. Who is this agent? What does it do? What does it NOT do? Set boundaries.

Example structure: > You are a {role description}. Your job is to {primary task}. You receive {input description} and produce {output description}. You do NOT {explicit boundary — e.g., "write final copy", "make API calls", "access external systems"}.

2. Context Section Describe what data the agent receives and in what format. This is where you document the pre-processed data structure so Claude knows what to expect.

Example structure: > ## Input Data > You receive a JSON array of items. Each item has these fields: > - title (string): ... > - url (string): ... > - score (number): ... > {list every field the stripped JSON contains}

3. Task Steps Numbered, specific, unambiguous steps. Each step should be one action. No compound steps.

Rules for writing task steps:

  • Start each step with a verb
  • Include the decision criteria inline (e.g., "Score each item from 1-10 based on: relevance to {topic} (0-3), recency (0-3), engagement (0-2), authority (0-2)")
  • Specify what to do with items that fail the filter (skip silently, log, or flag)
  • Reference tool names explicitly if the agent uses MCP tools or file I/O

4. Output Format Specify exactly what the agent should produce. Include a template or schema. If the output goes to a database, specify the column names and data types. If it's a file, specify the format.

5. Rules and Constraints Numbered list of non-negotiable rules. These are the guardrails. Common rules:

  • Never fabricate data
  • Never exceed {N} output items per run
  • Always deduplicate against {source} before writing
  • If unsure about a classification, choose {default}
  • Never mention {brand/product} directly in {context}

6. Example Output One complete, realistic example of the agent's output. Not abbreviated. The full thing, including edge cases. This is the single most effective way to steer Claude's behavior.

Present the system prompt outline (section summaries, not the full text yet) and ask the user to confirm the scope and constraints.


3C: Config File Design

Design the YAML configuration file. This file makes the agent portable and tunable without editing code.

agent:
  name: [agent-name]            # lowercase, hyphens, no spaces
  type: [pipeline|dual-layer|multi-agent|feedback-loop]
  schedule: [description]       # e.g., "Mon-Fri 09:00" or "every 2 days"
  model: [claude-sonnet-4-6|claude-opus-4-6]
  max_turns: [number]           # how many tool-use turns Claude gets
  allowed_tools: []             # list of MCP tools or built-in tools

data_sources:
  - name: [source-name]
    type: [api|scrape|file|database]
    endpoint: [URL or path]
    auth_env_var: [env var name] # e.g., APIFY_TOKEN
    rate_limit: [requests/min]
    fields_to_keep: []          # list of fields to preserve after stripping
    max_items: [number]         # cap per source

budget:
  max_per_run: [dollar amount]  # total budget ceiling per execution
  tracking_file: [path]         # e.g., logs/spend_{agent-name}.json
  alert_threshold: [percentage] # warn when spend hits this % of max

output:
  type: [file|database|api|notification]
  destination: [path or endpoint]
  format: [markdown|json|csv]
  max_output_items: [number]    # cap on items produced per run

filters:
  min_score: [number]           # minimum relevance score to include
  max_age_hours: [number]       # ignore items older than this
  dedup_field: [field name]     # field to use for deduplication
  dedup_source: [path or DB]   # where to check for existing items

schedule:
  type: [launchd|cron|manual]
  weekdays_only: [true|false]
  min_hours_between_runs: [number]  # gate mechanism
  timezone: [timezone string]       # e.g., "America/New_York"

Present the proposed config values to the user and ask for adjustments.


3D: Scheduling Design

Based on the user's OS and schedule requirements, design the scheduling configuration.

For macOS (launchd):

Generate a .plist file:


    Label
    com.{org}.{agent-name}
    ProgramArguments
    
        /path/to/run_{agent-name}.sh
    
    StartCalendarInterval
    
        
        
            Weekday1
            Hour9
            Minute0
        
        
    
    StandardOutPath
    /path/to/logs/{agent-name}.stdout.log
    StandardErrorPath
    /path/to/logs/{agent-name}.stderr.log
    EnvironmentVariables
    
        PATH
        /usr/local/bin:/usr/bin:/bin
    

Important macOS note: If the agent needs to access files in ~/Desktop/ or ~/Documents/, macOS TCC (Transparency, Consent, and Control) will block launchd-spawned processes. The fix is to use a compiled C launcher binary that has been granted Full Disk Access in System Settings. The launcher simply execv()s the real shell script, and FDA propagates to the child process.

For Linux (cron):

Generate a crontab entry:

# {agent-name} — runs Mon-Fri at 09:00
0 9 * * 1-5 /path/to/run_{agent-name}.sh >> /path/to/logs/{agent-name}.cron.log 2>&1

Gate mechanism (both platforms): The schedule triggers the script, but the script's own gate check (Step 3A) decides whether to actually run. This means the schedule can fire daily, but the gate ensures minimum spacing between actual executions. This is simpler and more reliable than trying to encode complex scheduling logic in cron or launchd.


Step 4 — Cost Estimation

Calculate the estimated cost per run and monthly projection. Present this as a table.

Token Estimation:

| Component | Calculation | |---|---| | System prompt | Word count x 1.3 = input tokens | | Input data | Items x avg bytes per item x 0.25 = input tokens (1 token ~ 4 bytes for English) | | Output | Expected output items x avg words per item x 1.3 = output tokens |

Claude API Pricing (current as of 2026):

| Model | Input (per 1M tokens) | Output (per 1M tokens) | |---|---|---| | claude-sonnet-4-6 | $3.00 | $15.00 | | claude-opus-4-6 | $15.00 | $75.00 |

Cost Per Run:

| Line Item | Estimate | |---|---| | Data source API costs | $ {based on user's APIs} | | Claude input tokens | $ {system prompt + data} | | Claude output tokens | $ {expected output} | | Total per run | $ {sum} | | Monthly projection | $ {total x runs per month} |

Present the estimate and ask: "Does this monthly cost work for your budget? If not, we can reduce cost by: (1) switching to Sonnet, (2) stripping more fields from the input data, (3) reducing the number of items per run, or (4) running less frequently."


Step 5 — Data Optimization Tips

Before writing files, present these optimization patterns. Highlight which ones are relevant to the user's specific agent.

1. JSON Stripping (always relevant) Remove unnecessary fields from API responses before sending to Claude. This is the single biggest lever for reducing cost. Write a small preprocessing script that keeps only the fields Claude needs.

Example: Raw Reddit post from the API might includ

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.