Install
$ agentstack add skill-shimo4228-contemplative-agent-replayable-audit-logs ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Replayable Audit Logs (observability by default)
> Sibling genre: an instrument reads across the whole store at query time > (distributions, calibration) — see > [read-only-instruments](../read-only-instruments/SKILL.md). The log is the > corpus; the instrument is one lens over it.
Principle (canonical rationale: [ADR-0075](../../../docs/adr/0075-observability-by-default.md)): a feature with external I/O, LLM calls, or heuristic decisions ships its audit log in the same PR — because the corpus must predate the failure it will one day explain. The exemplar loop: verification-audit.jsonl existed for weeks as an ordinary side effect → the round-6/7 parser repairs (ADR-0062) replayed 792 real challenges offline with a zero-wrong hard gate. Ad-hoc logging added at investigation time can never provide that.
Record schema checklist
Design the record so the run can be replayed offline, not merely read:
- [ ] Raw input, recoverable — the exact input the decision saw. Untrusted
text (API responses, user/content text) as base64 + sha256, never free text: a raw log read must not become a prompt-injection path. Bound the stored size (challenge_truncated-style flag when cut).
- [ ] Decision path — which branch/tier handled it (e.g. `solver_path:
codeparse | llmextract | llm_reason | none`).
- [ ] Reason codes, categorical — every abstain / fallback / failure gets a
machine-groupable code (abstain_reason: reasoning_self_inconsistent), not prose. A silent fallback is a defect, not a style choice.
- [ ] Outcome — what happened downstream (
verify_success), plus a
sanitized error (strip_to_printable, length-capped).
- [ ] Timestamps + stable keys —
ts(ISO, UTC) and a content hash
(sha256) so records dedupe and join across retries.
Writer: append-only JSONL under MOLTBOOK_HOME/logs/ via append_jsonl_restricted (restricted permissions, best-effort — the feature must not fail because logging failed).
Ground truth discipline
A log becomes a labeled corpus when outcomes are recorded honestly:
- Positive truth — an externally accepted result pins the correct answer
for that exact input (keyed by input hash).
- Negative truth — an externally rejected result is durably wrong for
that input, with no manual labeling. Rejections are data; log them with the same fidelity as successes.
- Manual labels — for inputs the outside world never confirmed, keep a
hand-labeled file next to the replay script (manual_labels.json), each label with provenance (hand-solved / twin-confirmed against an accepted same-shape input). A null answer marks known-unresolvable cases (external inconsistency) so the gate can excuse rather than ignore them.
Replay harness pattern
Copy the template in docs/evidence/adr-0062-parser-rewrite/:
- Pure-code script (no LLM, no network) that re-runs the deterministic layer
over every unique logged input.
- Hard gate: zero wrong vs known truth (positive + negative). Coverage is
a soft metric — report it, don't gate on it.
- Unlabeled parses are printed for labeling; the gate fails until they are
labeled or the behavior abstains.
- Exit code 0 only on gate PASS, so the harness can sit in a chain.
Corpus-driven repair loop
When a feature "keeps failing at the same rate":
- Aggregate before hypothesizing — success/failure per decision path over
the log; the failing tier is rarely the one you suspect.
- Decode and classify every failure into named classes.
- Twin-confirm intended fixes — before changing a rule, find an accepted
same-shape record proving the external system's semantics; a fix without a twin is a guess.
- Fix behind the replay hard gate (no lost previously-correct cases), and pin
the failure round as regression fixtures cut from real traffic (base64 in tests, with provenance comments).
- Record the round as an ADR amendment; leave irreducible cases explicitly
labeled (the floor is part of the finding).
Verify-gate question
At the chain's Verify step, for any feature in scope:
> この機能が誤動作したとき、どのログを見れば理由が分かるか。 > そのログからオフラインで再現(リプレイ)できるか。
No answer → back to design, in the same PR. Existing instruments to imitate: verification-audit.jsonl (solver), api-audit.jsonl (API drift), audit.jsonl (approval gates), LLM telemetry caller tags (retunes justified by measured done_reason rates).
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: shimo4228
- Source: shimo4228/contemplative-agent
- License: MIT
- Homepage: https://doi.org/10.5281/zenodo.19212118
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.