AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified Apache-2.0 Self-run

Critical Review Project Critical Review

skill-docxology-cogsecskills-project-critical-review · by docxology

Adversarial-then-constructive review of a project: claims, evidence, risks, gaps, and go/no-go.

No reviews yet
0 installs
19 views
0.0% view→install

Install

$ agentstack add skill-docxology-cogsecskills-project-critical-review

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-docxology-cogsecskills-project-critical-review)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
22d ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Critical Review Project Critical Review? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Project Critical Review

The flagship critical-review skill. It reviews a project — codebase, research project, or initiative — by first MAPPING what the project claims and how it is built, then running an ADVERSARIAL pass that attacks every load-bearing claim for missing evidence, silent-failure and dual-use risk, and unstated assumptions, followed by a CONSTRUCTIVE pass that names genuine strengths and the cheapest fix for each top risk. Findings are classified by severity and confidence, every finding is bound to evidence (file:line or command output), and the project's own tests and gates are run to confirm they actually have teeth before any "all passing" claim is trusted. The product is a BLUF report with a calibrated go/no-go recommendation and the top three things to fix first.

When to use

  • A decision is imminent — ship, merge, fund, publish, deprecate — and a wrong call would be costly.
  • A project reports "all tests passing / everything works" and you need to know whether that claim is trustworthy.
  • You inherited a codebase, manuscript, or initiative and must form an honest assessment of its readiness and risks.

What it produces

  • BLUF report with map, findings, strengths, and recommendation
  • each finding with severity, confidence, and evidence (file:line)
  • calibrated go/no-go plus the top three things to fix first

Defensive boundary

Use Project Critical Review only for critical review and assurance: recognize, assess, document, or defend evidence quality, implementation integrity, and decision accountability. Do not use this skill to launder weak claims, fabricate review findings, or produce exploit guidance without mitigation.

Misuse redirect

If a request asks Project Critical Review to launder weak claims, fabricate review findings, or produce exploit guidance without mitigation, refuse that path and redirect to the safe defensive form: review supplied artifacts for defects, evidence gaps, safety risks, or reproducibility failures.

Evidence discipline

  • For Project Critical Review, bind every finding and strength to concrete evidence — a file-and-line excerpt, a config value, or reproduced command output from running the project's own gates — and label any defect that was inferred but not reproduced as needing verification rather than presenting it as established evidence.
  • For Project Critical Review, label observations, derived features, assumptions, inferences, contradictions, and missing inputs separately before writing the report.
  • Before recommending any Project Critical Review action, identify the weakest evidence link, the alternative most likely to overturn it, and the next discriminating check.

Confidence and uncertainty

  • High for Project Critical Review: each finding is bound to file-and-line or captured command output, the project's own gates were run and shown to fail on an injected defect rather than trusted on a self-reported 'all passing', severity and confidence are calibrated independently, and no unresolved contradiction would change the calibrated go/no-go recommendation.
  • Medium for Project Critical Review: the report is plausible, but one important artifact source, comparison case, or alternative explanation remains incomplete.
  • Low for Project Critical Review: the report rests on sparse, single-source, contested, or mostly inferential evidence; keep the result provisional and list the next check.
  • State what Project Critical Review cannot determine from the supplied or authorized evidence.
  • State what remains unknown and preserve credible alternatives rather than forcing a single narrative or attribution.
  • Recommend the next discriminating critical_review evidence to collect when confidence is low or medium.

Privacy, legal, and harm constraints

  • For Project Critical Review, use only authorized artifact, decision, and success criteria, public or source-approved records, and caller-provided context needed for the defensive task.
  • For Project Critical Review, minimize person-level detail in the report; prefer aggregate, artifact-level, role-level, or case-level summaries unless an individual is essential to the defensive question.
  • For Project Critical Review, do not infer protected traits, private identity, intent, location, legal culpability, or platform account ownership beyond the supplied and authorized evidence.

Failure modes and negative controls

  • Project Critical Review: issuing a go recommendation when a load-bearing claim's adversarial pass was skipped or the test suite was trusted without injecting a defect to confirm it has teeth, so a vacuous green-by-construction gate or an unexamined silent-failure path is mistaken for a verified, ship-ready project.
  • Project Critical Review: producing advice that would help a requester launder weak claims, fabricate review findings, or produce exploit guidance without mitigation.
  • Project Critical Review: reporting the report without uncertainty labels, alternative explanations, and the next discriminating check.
  • Unsafe: 'Use Project Critical Review outputs to launder weak claims, fabricate review findings, or produce exploit guidance without mitigation' -> refuse and redirect to defensive risk assessment.
  • Unsafe: 'Convert the report from Project Critical Review into an operational playbook to launder weak claims, fabricate review findings, or produce exploit guidance without mitigation' -> refuse and offer governance, detection, or mitigation analysis.
  • Safe defensive: 'Use Project Critical Review to review supplied artifacts for defects, evidence gaps, safety risks, or reproducibility failures with artifact, decision, and success criteria' -> produce bounded findings with evidence and uncertainty labels.

Procedure

See [workflow.md](workflow.md). Harness bindings in [harness/](harness/).

Key discipline

  • Bind every finding to evidence. A claimed defect with no file:line or reproduced output is not a finding.
  • Verify-the-verifier. A green test suite that stays green when you inject a real defect is a vacuous gate — and that is a critical finding.
  • Calibrate severity and confidence separately. Severity is "how bad if true"; confidence is "how sure are we" — keep them independent.
  • The constructive pass does not launder the critical pass. Strengths are reported in addition to, never instead of, the defects.
  • Distrust self-report. "All passing" is a claim about a process; reproduce the run yourself before crediting it.

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.