AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Factor Benchmark

skill-minihellboy-factorminer-factor-benchmark · by minihellboy

Run FactorMiner benchmark workflows — the Table 1 Top-K freeze benchmark, memory and strategy ablations, transaction-cost pressure tests, and the full suite. Use to compare FactorMiner against baselines or to reproduce paper results. Triggers on "benchmark", "ablation", "compare to baseline", "reproduce table 1", "cost pressure", "benchmark suite".

No reviews yet
0 installs
10 views
0.0% view→install

Install

$ agentstack add skill-minihellboy-factorminer-factor-benchmark

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-minihellboy-factorminer-factor-benchmark)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Factor Benchmark? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Factor Benchmark

This skill runs FactorMiner's canonical benchmark surface — the rigorous comparison layer that turns a single run into evidence.

Modes

| Mode | What it answers | |---|---| | table1 | Top-K freeze benchmark across configured universes vs. baselines — the headline reproduction. | | ablation-memory | How much does experience memory contribute? | | ablation-strategy | Effect of memory policy × dependence metric × backend. | | cost-pressure | How does the library hold up under rising transaction costs? | | efficiency | Operator- and factor-level runtime/compute cost. | | suite | The full benchmark suite in one run. |

Workflow

Run a benchmark

factorminer -o output/bench benchmark table1 --data path/to/market_data.csv
factorminer -o output/bench benchmark suite --data path/to/market_data.csv

Pass a pre-mined library to benchmark a specific run rather than mining fresh:

factorminer -o output/bench benchmark table1 \
  --data market_data.csv \
  --factor-miner-library output/run1/factor_library.json

efficiency takes no data — it profiles the engine itself.

Read the result

The CLI prints a per-universe summary (library IC, ICIR, avg |ρ|) and writes JSON payloads into the output directory. Fold those JSON files into the research note with factor-report --benchmark.

Interpreting ablations

  • An ablation that removes a feature and barely moves the metric means that feature is not earning its compute on this dataset — report that plainly.
  • cost-pressure is the honesty check: a library that only wins at zero cost is not a result.

Guardrails

  • Benchmark numbers are comparative research evidence, not a performance guarantee.
  • Use the same dataset and splits across compared runs, or the comparison is meaningless.
  • Reproduction claims must cite the exact config and run directory.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.