Install
$ agentstack add skill-phanghonghao-thu-awesome-skills-unilab-rollout-trace ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
UniLab Rollout Trace
Record a headless numeric rollout from a UniLab off-policy checkpoint, plot it into a stability figure, and judge whether the policy is stable. General-purpose: any task, as long as it's SAC/FlashSAC on MuJoCo.
What it does
Given a checkpoint + task, produce three artifacts next to the checkpoint:
| Artifact | Content | |---|---| | rollout_trace_.npz | Full per-step tensors (dofpos/vel, actions, commands, basepos/quat, linvel, gyro, reward, rewardterms, gaitphase) | | rollout_trace_.summary.csv | Compact CSV (fixed schema) — the plotting input | | rollout_trace_.metadata.json | ctrldt, task, checkpointpath, jointnames, rewardlog_keys | | rollout_trace_.plot.png | 4-panel stability figure (base height, vel tracking, effort, reward) |
Then a one-line Chinese verdict: 稳定 / 不稳定 (fell at Xs).
Hard constraints (read first)
The recorder UniLab/scripts/record_offpolicy_rollout.py has these fixed limits — verify the target matches before promising anything:
| Dimension | Limit | Why | |---|---|---| | sim | MuJoCo only | sim_backend="mujoco" is hardcoded + uses mujoco.MjModel | | algo | SAC or FlashSAC only | uses create_sac_playback_session. PPO checkpoints CANNOT use this tool. | | task | any registered task | --task is a free arg; g1_walk_flat is only the default |
For PPO checkpoints, use the viser / MuJoCo-native viewer path from unilab-train instead — there is no numeric-trace recorder for PPO.
Where each step runs
| Step | Where | Why | |---|---|---| | Record (rollout) | WSL | needs UniLab env + mujoco | | Plot (CSV → PNG) | Windows python | has matplotlib 3.10 + numpy; CSV is already on D:\ |
The recorder only emits CSV/npz — plotting is done by this skill's own script scripts/plot_rollout_trace.py (numpy + matplotlib, no pandas).
Read first
references/recording.md — before running anything that depends on:
- The recorder's exact flags / return values
- CSV vs npz schema and which columns exist for non-locomotion tasks
- How to inject a specific command (zero / forward / random)
- Stability thresholds (first-terminated step, basez vs minbase_height)
- Advanced: npz reward-term decomposition
- Gotchas (commands/gait_phase fallback to zeros for non-locomotion tasks)
Treat references/recording.md as the source of truth for recorder specifics.
Slash-flag invocation
| Flag | Action | |---|---| | (none) or --record | Record a trace for a checkpoint, copy to Windows, plot, verdict | | --plot-only | Skip recording — just plot existing CSV(s) | | --compare | Overlay multiple CSVs in one figure (compare ckpts / cmds / tasks) | | --cmd zero\|fwd\| | Command to inject during recording (default zero) |
If the user gives a bare checkpoint path or task name, treat it as --record.
Full workflow (--record, default)
Inputs to resolve from the user's message:
TASK— e.g.g1_walk_flat(the task name, not a path)CKPT— absolute path to amodel_*.pt(WSL path if recording, or derive)CMD—zero(default) |fwd(~1 m/s forward) | explicitvx,vy,yaw
Step 1 — Record in WSL
Output goes next to the checkpoint so Windows plotting finds it directly. Pick a ` for the command type (zerocmd, fwdcmd, cmd10_0`, …).
wsl -d Ubuntu -u u20174 -- bash -c "export HOME=/home/u20174 && cd /home/u20174/UniLab && \
HF_ENDPOINT=https://hf-mirror.com /home/u20174/.local/uv run --no-sync \
python scripts/record_offpolicy_rollout.py \
--algo sac --task --sim mujoco \
--steps --num-envs 1 \
--output /mnt/d/Desktop_Files/UniLab/checkpoints//mujoco//rollout_trace_.npz \
algo.load_run=/mnt/d/Desktop_Files/UniLab/checkpoints//mujoco//model_.pt \
interactive.action_mode=policy \
"
- `
: 1000 is a good default for stability.1 step ≈ ctrl_dt s`
(G1 = 0.02 s → 1000 steps = 20 s).
--num-envs 1on CPU (more envs = slower, no benefit for a trace).- Command injection: the env reads commands internally; for a custom
command use a Hydra/env override appropriate to the task, e.g. +env.command_ranges_x='[0.0,0.0]' for zero, or +env.command_ranges_x='[1.0,1.0]' for fixed 1 m/s. If unsure, see references/recording.md — when in doubt, record zero-command (the strictest stability test).
The recorder prints summary_env=0 first_terminated_step=.
Step 2 — Plot on Windows python
The CSV is already on D:\ (next to the checkpoint). Run the skill's plotter:
python "C:/Users/20174/.claude/skills/unilab-rollout-trace/scripts/plot_rollout_trace.py" \
"D:/Desktop_Files/UniLab/checkpoints//mujoco//rollout_trace_.summary.csv" \
--min-base-height 0.3
--min-base-heightis optional but recommended for bipeds (draws the fall
threshold on the base-height panel). Get the value from the task yaml (min_base_height) or env cfg; common: G1 = 0.3.
- Auto-opens the PNG unless
--no-show. - Prints a compact summary table (steps, durations, fall time, minz,
mean/cum reward, effort).
Step 3 — Verdict
From the printed summary table:
| Signal | Stable | Unstable | |---|---|---| | fall@ | — (survived full run) | a time (terminated) | | min_z vs min_base_height | stays above | drops below | | mean_r | positive / near command target | strongly negative |
One-line Chinese verdict + next action, e.g.:
- 稳定: "零命令下站满 20s 未倒,base_z 稳在 0.72m,无需干预。"
- 不稳定: "零命令下 1.16s 倒下,base_z 跌到 0.12m " --min-base-height 0.3
## Response style
Order:
1. What was recorded (task / ckpt / command / steps)
2. The summary table (from the plotter stdout)
3. One-line Chinese verdict
4. Recommended next action (continue train / change ckpt / tune termination)
Keep concise. Offer the npz reward-term breakdown (`references/recording.md`)
only if the user wants to dig into *why* it's unstable.
## Common issues
| Symptom | Cause | Fix |
|---|---|---|
| `Key 'X' is not in struct` | new Hydra/env key | prefix with `+env.X=...` |
| `--load-run must be '-1' or dir name` | recorder Hydra override vs CLI | recorder already uses `algo.load_run=` override — keep it |
| PPO checkpoint | recorder is SAC/FlashSAC-only | use viser / MuJoCo viewer from `unilab-train` |
| Motrix target | recorder is MuJoCo-only | no headless trace for Motrix; use `eval --render-mode record` mp4 |
| `cmd_vx/yaw` columns all zero | non-locomotion task (no `info["commands"]`) | expected — those panels just show zeros |
| `ModuleNotFoundError: matplotlib` | plotting in WSL instead of Windows | run plotter on **Windows python** (it has matplotlib) |
| PNG won't open | non-Windows or `--no-show` set | open `D:/.../*.plot.png` manually |
## Reference files
- `references/recording.md` — recorder flags, CSV/npz schema, command injection,
stability thresholds, npz reward-term decomposition, gotchas.
- `scripts/plot_rollout_trace.py` — the plotter (numpy + matplotlib).
## Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- **Author:** [phanghonghao](https://github.com/phanghonghao)
- **Source:** [phanghonghao/THU-Awesome-Skills](https://github.com/phanghonghao/THU-Awesome-Skills)
- **License:** MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.