# Agent Research Aggregator

> Pre-pipeline aggregator that scans AI agent cache directories (.claude, .cursor, .antigravity, .openclaw) or any user-specified directory for experimentation logs, extracts insights and numeric results, and formats them as PaperOrchestra-ready inputs (idea.md + experimental_log.md). TRIGGER when the user says "aggregate my agent logs for paper writing", "extract experiments from my coding agent h…

- **Type:** Skill
- **Install:** `agentstack add skill-woodfishhhh-ez-math-model-agent-research-aggregator`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [woodfishhhh](https://agentstack.voostack.com/s/woodfishhhh)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [woodfishhhh](https://github.com/woodfishhhh)
- **Source:** https://github.com/woodfishhhh/EZ_math_model/tree/main/skills/ez-math-model/external/paper-orchestra/skills/agent-research-aggregator

## Install

```sh
agentstack add skill-woodfishhhh-ez-math-model-agent-research-aggregator
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# agent-research-aggregator

---

## Should I run? (decision gate)

Before starting Phase 1, check whether aggregation is actually needed:

| Situation | Action |
|---|---|
| `workspace/inputs/idea.md` **and** `workspace/inputs/experimental_log.md` both exist and are non-empty | **Skip this skill entirely.** Proceed directly to `paper-orchestra`. |
| Either file is missing or empty, **and** the user provided a directory path | **Run this skill** with that directory as `--search-roots`. |
| Either file is missing or empty, **and** no directory was provided | Scan cwd and `~` by default; show the discovery summary to the user before continuing. |
| The inputs exist but look thin (e.g. idea.md has  \
    --agents  \
    --depth  \
    --since  \
    --out workspace/ara/discovered_logs.json
```

The script exits with code **2** when no `--project` filter is set (this is
expected on the first run). It prints a **"Projects found"** list to stdout —
show it to the user immediately.

**If no logs are found at all:** stop and ask the user to specify
`--search-roots` or point you at a directory that contains agent cache folders.

---

## Phase 1.5 — Project Selection (mandatory)

**A paper can only be written from a single project. You must ask the user
which project to use before any LLM processing begins.**

1. Display the numbered project list from the discovery summary, e.g.:
   ```
   Projects found:
     [1] /home/alice/projects/my-rl-experiment  (42 files)
     [2] /home/alice/projects/llm-eval-suite    (17 files)
     [3] /home/alice/projects/old-demo          (3 files)
   ```
2. Ask: *"Which project should this paper be based on? Please choose a number
   or paste the project path."*
3. **Do not proceed to Phase 2 until the user has answered.**
4. Re-run discovery with the chosen project to filter the manifest:

```bash
python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots  \
    --agents  \
    --depth  \
    --since  \
    --project "" \
    --out workspace/ara/discovered_logs.json
```

This overwrites `discovered_logs.json` so only the selected project's files
remain. The script exits 0 on success.

**If the discovery finds only one project:** skip the question and inform the
user: *"Only one project found: ``. Using it for the paper."* — then
re-run with `--project` automatically.

**If the discovery summary shows irrelevant files after filtering:** ask the
user whether to include or exclude them before continuing to Phase 2. Err on
the side of inclusion — the extraction prompt is conservative.

---

## Phase 2 — Extraction (LLM-assisted)

Process discovered logs in **batches** (group by agent type; keep batches under
~50 KB of raw text to stay within context limits):

For each batch:

1. **Read** the log files in the batch (the script's `--list` output tells you
   which file paths to read).
2. **Apply the extraction prompt** from `references/extraction-prompt.md` as
   your system message.
3. **Pass the raw log text** as the user message.
4. **Collect the structured JSON** the LLM returns (see schema in the prompt).
5. **Append** to `workspace/ara/raw_experiments.json`.

After all batches:

```bash
python skills/agent-research-aggregator/scripts/extract_experiments.py \
    --discovered workspace/ara/discovered_logs.json \
    --out workspace/ara/raw_experiments.json \
    --validate-only
```

Run this in `--validate-only` mode to check the combined JSON is well-formed
and meets the minimum schema (`experiments` array non-empty, each entry has
`hypothesis` or `method` or `results`). Fix any malformed entries before Phase 3.

---

## Phase 3 — Synthesis (LLM-assisted)

Consolidate possibly-redundant experiment records from multiple agent caches into
a single coherent research narrative. This is ONE LLM call.

**System message:** Use `references/synthesis-prompt.md` verbatim.

**User message:**
```

{contents of workspace/ara/raw_experiments.json}

```

The LLM must return a `synthesis.json` with keys:
- `research_question` — the overarching question being investigated
- `hypothesis` — the core proposed solution / claim
- `method_summary` — how the approach works (concise, no data leakage)
- `key_contributions` — 2–5 bullet strings
- `experimental_setup` — datasets, metrics, baselines, implementation notes
- `results_tables` — array of `{title, headers[], rows[]}` markdown-table objects
- `qualitative_observations` — free-form text blocks (what worked, what didn't,
  failure modes, ablation insights)
- `iteration_history` — ordered list of `{iteration_id, change_description,
  outcome}` entries if multiple iterations are detected
- `open_questions` — questions that remain unanswered in the logs

Save to `workspace/ara/synthesis.json`.

> **Note:** By this point, the user has already selected a single project in
> Phase 1.5. The synthesis should represent one coherent research thread. If
> the LLM still surfaces multiple disconnected research questions, flag this
> as a data quality warning in the audit report (Phase 5) but do not re-ask
> for project selection — that decision was made earlier.

---

## Phase 4 — Formatting (deterministic)

Convert `synthesis.json` into PaperOrchestra input files:

```bash
python skills/agent-research-aggregator/scripts/format_po_inputs.py \
    --synthesis workspace/ara/synthesis.json \
    --out workspace/inputs/
```

This generates two files:

### `workspace/inputs/idea.md` (Sparse variant)

Follows the PaperOrchestra Sparse Idea format (arXiv:2604.05018, §3.1):

```markdown
# [Synthesized Research Title]

## Problem

## Hypothesis

## Method

## Key Contributions

## Open Questions

```

### `workspace/inputs/experimental_log.md`

Follows the PaperOrchestra Experimental Log format (App. D.3):

```markdown
## 1. Experimental Setup

## 2. Raw Numeric Data

## 3. Qualitative Observations

### Iteration History

```

After running the script, **review both files** with the user:

1. Read `workspace/inputs/idea.md` aloud and ask: "Does this accurately capture
   your research question and method?"
2. Read the table headers from `workspace/inputs/experimental_log.md` and ask:
   "Are these the correct metrics and baselines?"

Revise based on feedback before proceeding to PaperOrchestra.

---

## Phase 5 — Audit Report (deterministic)

```bash
python skills/agent-research-aggregator/scripts/format_po_inputs.py \
    --synthesis workspace/ara/synthesis.json \
    --out workspace/inputs/ \
    --report workspace/ara/aggregation_report.md
```

The `--report` flag makes the script also write `aggregation_report.md`, which
contains:

- Number of agent caches scanned, files read, batches processed
- Per-agent breakdown (files found per agent type)
- Experiment records extracted (count, date range)
- Iterations detected (count, convergence direction)
- Data quality warnings (gaps, low-confidence extractions, conflicting numbers)
- Files written and their sizes

Show the report to the user. If the data quality section lists warnings, discuss
them before running paper-orchestra — garbage in, garbage out.

---

## Handoff to PaperOrchestra

Once the user has confirmed `idea.md` and `experimental_log.md`, the workspace
is ready for the paper-orchestra pipeline. You still need:

| File | Status | Action |
|---|---|---|
| `workspace/inputs/idea.md` | ✓ generated | user review recommended |
| `workspace/inputs/experimental_log.md` | ✓ generated | user review recommended |
| `workspace/inputs/template.tex` | **MISSING** | ask user to provide their conference LaTeX template |
| `workspace/inputs/conference_guidelines.md` | **MISSING** | ask user to provide (page limit, deadline, formatting rules) |

Tell the user exactly which two files are still needed, then offer to run
`paper-orchestra` once they supply them.

---

## Error handling

| Situation | Action |
|---|---|
| Cache directory does not exist | Skip silently; note in report |
| File is binary or non-text | Skip; note in report |
| File > 200 KB | Truncate at 200 KB; note in report with path |
| LLM extraction returns malformed JSON | Re-prompt once with the parse error appended; if still malformed, log the batch as `status: failed` and continue |
| Synthesis returns > 1 `research_question` | Log as data quality warning in audit report; do not re-ask for project (was selected in Phase 1.5) |
| `results_tables` is empty after synthesis | Warn the user — PaperOrchestra's section-writing agent needs numeric data |

---

## Hard rules (never violate)

1. **Never write to agent cache directories.** This skill is read-only on `.claude/`, `.cursor/`, `.antigravity/`, `.openclaw/`.
2. **Never include personal information** (emails, names, credentials, API keys) in generated `idea.md` or `experimental_log.md`. The extraction prompt instructs the LLM to strip PII; double-check before handoff.
3. **Never fabricate results.** If a metric appears in only one log with low confidence, mark it `[UNVERIFIED]` in the table rather than silently including it.
4. **Never proceed past Phase 1 without user confirmation** of the discovered file list if the scan found > 50 files.

---

## Quick reference

```bash
# Phase 1: discover all projects (exits with code 2 — project selection required)
python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots . ~ --out workspace/ara/discovered_logs.json

# Phase 1.5: re-run with chosen project (exits 0)
python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots . ~ \
    --project "/home/user/projects/my-chosen-project" \
    --out workspace/ara/discovered_logs.json

# ... (Phase 2: LLM extraction calls, see above) ...

python skills/agent-research-aggregator/scripts/extract_experiments.py \
    --discovered workspace/ara/discovered_logs.json \
    --out workspace/ara/raw_experiments.json --validate-only

# ... (Phase 3: LLM synthesis call, see above) ...

python skills/agent-research-aggregator/scripts/format_po_inputs.py \
    --synthesis workspace/ara/synthesis.json \
    --out workspace/inputs/ \
    --report workspace/ara/aggregation_report.md
```

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [woodfishhhh](https://github.com/woodfishhhh)
- **Source:** [woodfishhhh/EZ_math_model](https://github.com/woodfishhhh/EZ_math_model)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-woodfishhhh-ez-math-model-agent-research-aggregator
- Seller: https://agentstack.voostack.com/s/woodfishhhh
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
