AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified Apache-2.0 Self-run

T Demo Run

skill-timzaak-web-dev-skills-t-demo-run · by timzaak

Run a single demo E2E test file, diagnose failures, dispatch fixes to agents, and re-run until pass.

No reviews yet
0 installs
40 views
0.0% view→install

Install

$ agentstack add skill-timzaak-web-dev-skills-t-demo-run

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-timzaak-web-dev-skills-t-demo-run)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of T Demo Run? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

单文件 Demo 测试运行与修复

运行时边界统一参考:${CLAUDE_PLUGIN_ROOT}/protocols/runtime-boundaries.md

目标

  • 先对一个测试文件整体执行。
  • 整体失败时,再按用例粒度顺序执行。
  • 单个用例失败时先诊断,再分发到对应 agent 修复。
  • 修复后必须执行相关后端/前端补测,不能只跑 Demo。
  • 输出可恢复的任务状态与机器可解析结果。

使用方式

/t-demo-run demo/e2e/super-admin/super-admin-comprehensive-demo.e2e.ts

执行流程

  • 参数校验。
  • 测试文件必须存在且扩展名为 .e2e.ts
  • 运行前清理。
uv run scripts/cleanup-demo.py
  • 先运行整个测试文件。
uv run scripts/demo-test-runner.py "[测试文件]" --run-id [RUN_ID]
  • 若整个测试文件通过:
  • 不再拆分用例运行。
  • 直接输出结果 JSON。
  • 若整个测试文件失败,列出测试用例。
uv run scripts/demo-test-runner.py "[测试文件]" --list-tests
  • 为每个用例创建任务并顺序执行。
uv run scripts/demo-test-runner.py "[测试文件]" --run-id [RUN_ID] --grep "[测试标题]"
  • 单用例失败修复循环(最多 6 次)。
  • 先通过 Agent(subagent_type="demo-diagnose") 启动诊断 subagent,传入 testFile、runId、testCaseTitle,生成结构化诊断。
  • 按诊断结果通过 Agent tool 分发到对应修复 subagent:Agent(subagent_type="demo-dev") / Agent(subagent_type="frontend-dev") / Agent(subagent_type="backend-dev") / Agent(subagent_type="miniapp-dev")
  • 读取修复 agent 返回的 tests_to_run(必填)并校验字段:
  • layer: backend|frontend|miniapp|demo
  • command: 可直接执行命令
  • reason: 关联说明
  • required: 是否必须通过(默认 true
  • 执行补测(按层顺序串行):backend -> frontend -> miniapp -> demo
  • 补测命令必须来自允许入口:
  • 后端:uv run scripts/backend-test.py -- [filter]
  • 前端:cd frontend && npm run test:run -- [pattern]
  • 小程序:cd miniapp && npm run typecheckcd miniapp && npm run build:weapp
  • Demo:uv run scripts/demo-test-runner.py "[测试文件]" --run-id [RUN_ID] --grep "[测试标题]"
  • miniapp 补测只在目标项目存在 miniapp/ 或诊断报告明确归因到小程序交付线时执行;未启用 miniapp 的项目跳过该层。
  • 若 agent 未返回 tests_to_run
  • 记录契约缺失(P1)
  • 执行最小兜底补测(按改动层至少 1 条 backend/frontend/miniapp 相关测试)
  • 重新运行当前用例验证修复(即 demo 层验证)。
  • 结果输出。
  • 最后一行必须输出机器可解析 JSON,格式固定为:
Result: {"success":"true|false","fixed":"true|false","logs":"demo/test-results/runs/[RUN_ID]","exit_code":0,"test_file":"demo/e2e/...","run_id":"[RUN_ID]","error":""}
  • 字段含义:
  • success: 最终 Demo 验证是否通过。
  • fixed: 本次是否经历失败后修复并通过;首次整体通过时为 false
  • logs: 本次运行主日志目录,优先使用 demo-test-runner.py 返回的 logs
  • exit_code: 最终 Demo 验证退出码。
  • test_file: 输入测试文件路径。
  • run_id: 本次运行 ID。
  • error: 失败时的简要错误;成功时为空字符串。
  • 不再额外要求自然语言“最终总结”;必要说明只保留为失败诊断、修复记录或 error 字段。

恢复机制

当流程中断时:

  • 读取 TaskList
  • 找到 pending 且依赖已满足的任务继续执行。

失败处理

  • 环境启动失败:停止并记录错误。
  • 无可用修复方案:标记该用例失败,继续下一个。
  • 达到最大重试次数:标记失败并继续。
  • 补测失败:记录失败与风险,不阻断本用例修复循环,继续 Demo 重跑与后续尝试。

质量门禁

  • 单次执行只处理一个测试文件。
  • 必须先整体运行测试文件;只有整体失败时才拆分用例。
  • 拆分后的用例执行必须串行。
  • 每个失败用例必须有诊断记录。
  • 每次修复后必须先执行相关层补测,再执行 Demo 验证。
  • 必须输出最后一行 Result: {...}

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.