Install
$ agentstack add skill-heshamfs-materials-simulation-skills-simulation-validator Open-source listing — not yet scanned by AgentStack. Follow the source repository for install instructions.
Security review
⚠ Flagged1 finding(s); flagged for manual review. · v0.1.0 How review works →
- • Prompt-injection patterns
- • Secret / credential exfiltration
- • Dangerous shell & filesystem operations
- • Untrusted network calls
- • Known-malicious package signatures
- high Dangerous shell/eval execution.
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ● Shell / process execution Used
- ✓ Environment & secrets No
- ● Dynamic code execution Used
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Simulation Validator
Goal
Provide a three-stage validation protocol: pre-flight checks, runtime monitoring, and post-flight validation for materials simulations.
Requirements
- Python 3.10+
- No external dependencies (uses Python standard library only)
- Works on Linux, macOS, and Windows
Inputs to Gather
Before running validation scripts, collect from the user:
| Input | Description | Example | |-------|-------------|---------| | Config file | Simulation configuration (JSON/YAML) | simulation.json | | Log file | Runtime output log | simulation.log | | Metrics file | Post-run metrics (JSON) | results.json | | Required params | Parameters that must exist | dt,dx,kappa | | Valid ranges | Parameter bounds | dt:1e-6:1e-2 |
Decision Guidance
When to Run Each Stage
Is simulation about to start?
├── YES → Run Stage 1: preflight_checker.py
│ └── BLOCK status? → Fix issues, do NOT run simulation
│ └── WARN status? → Review warnings, document if accepted
│ └── PASS status? → Proceed to run simulation
│
Is simulation running?
├── YES → Run Stage 2: runtime_monitor.py (periodically)
│ └── Alerts? → Consider stopping, check parameters
│
Has simulation finished?
├── YES → Run Stage 3: result_validator.py
│ └── Failed checks? → Do NOT use results
│ → Run failure_diagnoser.py
│ └── All passed? → Results are valid
Choosing Validation Thresholds
| Metric | Conservative | Standard | Relaxed | |--------|--------------|----------|---------| | Mass tolerance | 1e-6 | 1e-3 | 1e-2 | | Residual growth | 2x | 10x | 100x | | dt reduction | 10x | 100x | 1000x |
Script Outputs (JSON Fields)
| Script | Output Fields | |--------|---------------| | scripts/preflight_checker.py | report.status, report.blockers, report.warnings | | scripts/runtime_monitor.py | alerts, residual_stats, dt_stats (alerts include NaN/Inf/overflow detection, residual growth, and dt collapse) | | scripts/result_validator.py | checks, confidence_score, failed_checks, status (PASS / FAIL / INSUFFICIENT_DATA); confidence_score is null when no check ran | | scripts/failure_diagnoser.py | probable_causes, recommended_fixes |
Three-Stage Validation Protocol
Stage 1: Pre-flight (Before Simulation)
- Run
scripts/preflight_checker.py --config simulation.json - BLOCK status: Stop immediately, fix all blocker issues
- WARN status: Review warnings, document accepted risks
- PASS status: Proceed to simulation
> Note: preflight_checker.py validates required keys, numeric ranges, > output-directory access, and disk space. It does not evaluate numerical > stability (CFL / diffusion-Fourier). For explicit stability gating use > skills/core-numerical/numerical-stability/scripts/cfl_checker.py.
python3 scripts/preflight_checker.py \
--config simulation.json \
--required dt,dx,kappa \
--ranges "dt:1e-6:1e-2,dx:1e-4:1e-1" \
--min-free-gb 1.0 \
--json
Stage 2: Runtime (During Simulation)
- Run
scripts/runtime_monitor.py --log simulation.logperiodically - Configure alert thresholds based on problem type
- Stop simulation if critical alerts appear
python3 scripts/runtime_monitor.py \
--log simulation.log \
--residual-growth 10.0 \
--dt-drop 100.0 \
--json
Stage 3: Post-flight (After Simulation)
- Run
scripts/result_validator.py --metrics results.json - All checks PASS: Results are valid for analysis
- Any check FAIL: Do NOT use results, diagnose failure
python3 scripts/result_validator.py \
--metrics results.json \
--bound-min 0.0 \
--bound-max 1.0 \
--mass-tol 1e-3 \
--json
For variational / gradient-flow models (Allen-Cahn, Cahn-Hilliard), add --variational to enforce a strict monotone non-increasing energy check.
Failure Diagnosis
When validation fails:
python3 scripts/failure_diagnoser.py --log simulation.log --json
Conversational Workflow Example
User: My phase field simulation crashed after 1000 steps. Can you help me figure out why?
Agent workflow:
- First, check the log for obvious errors:
``bash python3 scripts/failure_diagnoser.py --log simulation.log --json ``
- If diagnosis suggests numerical blow-up, check runtime stats:
``bash python3 scripts/runtime_monitor.py --log simulation.log --json ``
- Recommend fixes based on findings:
- If residual grew rapidly → reduce time step
- If dt collapsed → check stability conditions
- If NaN detected → check initial conditions
Error Handling
| Error | Cause | Resolution | |-------|-------|------------| | Config not found | File path invalid | Verify config path exists | | Non-numeric value | Parameter is not a number | Fix config file format | | out of range | Parameter outside bounds | Adjust parameter or bounds | | Output directory not writable | Permission issue | Check directory permissions | | Insufficient disk space at | Disk nearly full on the output volume | Free up space or reduce output | | Invalid parameter name | --required name has disallowed characters | Use only letters, digits, _, ., - | | range max ... must be greater than min | Inverted/degenerate --ranges or bounds | Ensure max > min | | must be a finite positive number | nan/inf/negative threshold supplied | Pass a finite positive value | | Log file too large | Log exceeds the 500 MB parse cap | Truncate or pre-filter the log |
Interpretation Guidance
Status Meanings
| Status | Meaning | Action | |--------|---------|--------| | PASS | All checks passed | Proceed with confidence | | WARN | Non-critical issues found | Review and document | | BLOCK | Critical issues found | Must fix before proceeding |
Confidence Score Interpretation
| Score | Meaning | |-------|---------| | 1.0 | All validation checks passed → proceed with confidence | | 0.75+ | Most checks passed, minor issues | | 0.5-0.75 | Significant issues, review carefully | | min` required
--min-free-gbis validated as a finite positive number (negatives, zero,nan,infrejected)--residual-growthand--dt-dropthresholds are validated as finite positive numbers--bound-minand--bound-maxare validated as finite numbers (nan/infrejected), and--bound-max > --bound-minis enforced;--mass-tolis validated as a finite positive number- Invalid input exits with code 2 and an explanatory message
File Access
preflight_checker.pyreads a single user-specified config file (JSON/YAML) and checks disk space on the volume hosting the resolved output directoryruntime_monitor.pyreads a single log file specified by--log; log files are size-limited (500 MB max) and rejected before parsing if largerresult_validator.pyreads a single metrics file (JSON) specified by--metricsfailure_diagnoser.pyreads a single log file specified by--log; log files are size-limited (500 MB max) before parsing- No scripts write to the filesystem; all output goes to stdout
Tool Restrictions
- Read: Used to inspect script source, references, config files, and simulation logs
- Bash: Used to execute the four Python validation scripts (
preflight_checker.py,runtime_monitor.py,result_validator.py,failure_diagnoser.py) with explicit argument lists - Write: Used to save validation reports; writes are scoped to the user's working directory
- Grep/Glob: Used to locate log files, config files, and search references
Safety Measures
- No
eval(),exec(), or dynamic code generation - All subprocess calls use explicit argument lists (no
shell=True) failure_diagnoser.pyuses hardcoded, pre-compiled diagnostic regex patterns;runtime_monitor.pyaccepts optional--residual-pattern/--dt-patternoverrides that are compiled withre.compile(noeval) and applied only to the user's own log- Diagnostic strings emitted in output are drawn from the skill's fixed cause/fix table, not interpolated from raw log content
Limitations
- Not a real-time monitor: Scripts analyze logs after-the-fact
- Regex-based: Log parsing depends on pattern matching; may miss unusual formats
- No automatic fixes: Scripts diagnose but don't modify simulations
References
references/validation_protocol.md- Detailed checklist and criteriareferences/log_patterns.md- Common failure signatures and regex patterns
Version History
- v1.2.2 (2026-06-24): Added a Verification checklist (evidence-based, tied to the four scripts' JSON outputs) and a Common pitfalls & rationalizations table to harden agent interpretation of validation verdicts.
- v1.2.0 (2026-06-23): Corrected diagnostic regexes (no false convergence/blow-up on healthy logs), direction-aware dt-collapse detection, NaN/Inf scan in runtime monitor, strict variational energy check, non-vacuous bounds/confidence, config-relative output-dir + correct-volume disk check, and implemented the documented input-validation/file-size safeguards
- v1.1.0 (2024-12-24): Enhanced documentation, decision guidance, Windows compatibility
- v1.0.0: Initial release with 4 validation scripts
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: HeshamFS
- Source: HeshamFS/materials-simulation-skills
- License: Apache-2.0
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.