Install
$ agentstack add skill-graph-robots-open-robot-skills-perceiving-objects ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
perceiving-objects
Single-path perception: detect → disambiguate (pairwise crop tournament) → segment → fuse to 3D → extract OBB. Each DINO detection is cropped and upscaled, and the target is found via binary "A or B?" comparisons of crop pairs — far more reliable on small targets than a one-shot Set-of-Marks letter pick (~30% → 97% on the LIBERO-PosVar object-ID study).
Multi-camera handling uses a safe wrist-fallback gate (not blind KD-tree fusion): identification defaults to the exterior view; the wrist (eye-in-hand) view is consulted ONLY when the exterior pick fails its own close-up verify AND the wrist pick passes its own. On the 4-suite / 200-frame regression study this was the only zero-regression policy (+2.5% net, 0/189 frames regressed; blind fuse/verify→wrist/ wrist-only all regressed). See perceive_dino_vlm.run docstring.
On the verified-exterior path the wrist views still contribute cloud geometry (never identity): wrist clouds of the same object — gated by the multiview intersection check, with a geometry-seeded SAM fallback (exterior cloud projected into the wrist frame) — are fused into the output cloud so the OBB recovers the top face / far side a single front view misses. A lone front view yields a sliver OBB biased toward the camera by half the object depth, and that off-centre pinch is the measured slip-during-transport failure mode on tall bottles/cartons.
When to use
- Uncluttered scenes with visually distinct targets.
- Platforms where only DINO + VLM + SAM3 + geometry are deployed.
- The default single-target 3D perception skill.
When NOT to use
- Cluttered scenes with many similar nearby distractors: strengthen the
pairwise tournament by passing object_description shape/appearance hints (see the note above) rather than relying on the bare label alone.
- Clean-all-items / multi-item loops that need a clean "no match" loop
terminator. Prefer perceiving-objects-oneshot.
Recommended subgraph state flow
3 states:
observe → perceive → filter_obb
State details:
> About object_name below: it is a literal Python string — the natural > noun phrase for the object you are perceiving, drawn from this subgraph's > description (e.g. "alphabet soup can", "basket", "red bowl"). It is > a constant per subgraph instance, NOT a binding. DO NOT write > Ref("in.object_name") or any other Ref(...); the coordinator does > not declare object_name as a subgraph input. Write the string directly, > e.g. "object_name": "basket". > > About object_description (wire it whenever the task gives hints): > also a literal Python string. When the task/workflow description carries > shape or appearance hints for the target (e.g. an "Object context" block > with shape_hint / expected_label, or adjectives in the instruction), > pass them through verbatim, e.g. > "object_description": "small rectangular box, blue and white packaging, > ~5 cm wide". The description is injected into BOTH the pairwise > tournament prompt ("It looks like: …") and the close-up verification > question ("It should look like: …"). This is what disambiguates > look-alike packaging (several LIBERO grocery items are small blue/white > boxes) and keeps the verify gate from rejecting a correct pick whose > rendered asset reads as a generic box — a rejection forces the > wrist-camera fallback, whose single top-down view degrades the OBB > height and downstream grasps. Omit it (default "") only when the task > provides no hints.
observe—type: tool,tool: "robot.get_observation",
inputs: {}. Connector tool; flat name only.
perceive—type: script, filescripts//perceive_dino_vlm.py
from this bundle. Inputs: cameras=Ref("observe.cameras"), object_name="basket" (replace with the actual target noun phrase from this subgraph's description), object_description="..." (the task's shape/appearance hints — see the note above; strongly recommended whenever hints exist), plus any optional fields (dino_prompt, etc.). Returns {found, cloud, mask, score}.
filter_obb—type: tool,
tool: "geometry.filter_and_compute_obb", inputs={"points": Ref("perceive.cloud")}. Returns {"obb": }.
Wiring the exit (HARD)
Use the linear edge filter_obb → found → END. When the target isn't found, perceive returns an empty cloud and filter_obb raises on it, so the subgraph's on_error: "not_found" catches that path automatically. Do NOT add any conditional edges on perceive — the linear path plus set_on_error is sufficient.
✅ Correct (the literal gap.builder calls you should emit):
sg.add_node("perceive", type="script",
script="scripts//perceive_dino_vlm.py",
inputs={"cameras": Ref("observe.cameras"),
"object_name": "small blue and white cream cheese",
# from the task's Object context / shape_hint block:
"object_description": ("small rectangular box, blue "
"and white packaging, ~5 cm wide")})
sg.add_node("filter_obb", type="tool",
tool="geometry.filter_and_compute_obb",
inputs={"points": Ref("perceive.cloud")})
# add_exit() creates the success-marker noop node AND registers the
# exit value. Do NOT also call sg.add_node("found", type="noop") — that
# would conflict with the node add_exit created.
sg.add_exit("found")
sg.add_edge("perceive", "filter_obb")
sg.add_edge("filter_obb", "found")
sg.add_edge("found", END)
sg.set_on_error("not_found")
Bind the subgraph outputs (ALL THREE — required, no exceptions):
sg.set_outputs(
target_obb=Ref("filter_obb.obb"),
target_mask=Ref("perceive.mask"),
target_cloud=Ref("perceive.cloud"),
)
(Replace target_* with this subgraph's actual name prefix — e.g. container_obb, container_mask, container_cloud when authoring the container subgraph.)
All three bindings walk into a field of the producing node's output dict: geometry.filter_and_compute_obb returns {"obb": ...} (bind Ref("filter_obb.obb"), NOT a bare Ref("filter_obb")), while _mask and _cloud walk into fields of perceive's output dict. perceive_dino_vlm.py already produces all three; emitting them unconditionally lets downstream subgraphs that need any of them (e.g. learned-grasp skills require _cloud) wire up without you having to anticipate which skill they'll use.
Hard rules
- Subgraph-level outputs MUST emit ALL THREE:
_obb,
_mask, AND _cloud. The cloud is the fused world-frame point cloud needed by learned-grasp skills; emit it unconditionally so the downstream agent can wire it without round-tripping. See references/perception_pipeline_invariants.md.
geometry.filter_and_compute_obbreturns{"obb": OrientedBoundingBox};
bind via Ref("filter_obb.obb") (walk into the obb field). See references/geometry_calling_conventions.md.
Required end states
| End state | Meaning | |---|---| | found | OBB + mask bound; route to next subgraph (typically a grasp skill). | | not_found | Route to abort (or to done in clean-all-items loops). |
See also
prompts/vlm_pairwise.md— the pairwise-tournament VLM prompt template.scripts/perceive_dino_vlm.py— the canonical perception script.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: graph-robots
- Source: graph-robots/open-robot-skills
- License: Apache-2.0
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.