# Flows Design Review

> >-

- **Type:** Skill
- **Install:** `agentstack add skill-cognitedata-builder-skills-flows-design-review`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [cognitedata](https://agentstack.voostack.com/s/cognitedata)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** Apache-2.0
- **Upstream author:** [cognitedata](https://github.com/cognitedata)
- **Source:** https://github.com/cognitedata/builder-skills/tree/main/skills/flows-design-review

## Install

```sh
agentstack add skill-cognitedata-builder-skills-flows-design-review
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Flows Design Review

This is **step 3** of the Flows app certification flow:

```
flows-app-brief  →  build  →  flows-code-review  →  flows-design-review (this skill)  →  flows-external-app-submit
```

This is the **manual design quality assessment** described in
[docs.cognite.com/cdf/flows/guides/quality-guidelines](https://docs.cognite.com/cdf/flows/guides/quality-guidelines).
Target overall average: **3.8 or higher** to be launch-ready.

## Operating rules

- **Automate first, ask second.** For every question Q1–Q10, run the probes listed below to gather hard evidence from the repo and **propose a draft score (1–5) with rationale** *before* asking the user. The user's job is to confirm or override the proposed score, not to grade from scratch. This dramatically reduces the manual burden.
- The **task walkthrough (Step 2)** is the one part that cannot be skipped — automation cannot tell whether a user "gets lost" navigating a screen. Capture it manually and use it to override the auto-derived scores where lived experience disagrees.
- Use `AskQuestion` for every score so answers are structured. For each question present three options: *(a) accept the draft score*, *(b) override with a specific score*, *(c) override + add a note*.
- Pre-fill user, tasks, and persona context from `App-Brief.md` frontmatter when present.

## Step 0 — Pre-scan before prompting

**Always pre-scan before asking the user anything.** Read these sources silently and surface what you found as *evidence* — never as scores, never auto-saved:

| Source | Use it for |
| --- | --- |
| `App-Brief.md` frontmatter | Pre-fill primary user (`userRole`), tasks (`oneSentenceStory`), success criteria |
| `package.json` | Confirm `@cognite/aura` is installed and surface its version (informs Q1) |
| Latest `reviews/code-review/feedback-round-/code-review-report.md` | Pull design-adjacent findings (accessibility, error handling, UX copy) and present them as evidence under Q4/Q10 |
| `src/**/*.{ts,tsx,css}` | Q1 probe — grep for hard-coded hex/rgb colors and raw `px`/`rem` values outside Aura tokens |
| `src/**/*.{ts,tsx}` | Q5 probe — `onClick` on non-button elements without `role`/`tabIndex` |
| `src/**/*.{ts,tsx}` | Q10 probe — icon buttons missing `aria-label`, `` without `alt`, missing focus styles |

Show the user the pre-scan results in your opening message before any scoring. They are starting points, not verdicts. The manual task walkthrough (Step 2) and user-assigned scores remain authoritative.

## Step 0b — Choose feedback round

Look at `reviews/design-review/`. If it doesn't exist, this is round 1. Otherwise increment to the next missing `feedback-round-/` directory.

## Step 1 — Confirm user and tasks

Per the docs, "the quality assessment is only as useful as the clarity of the user and tasks it's based on."

If `App-Brief.md` exists, parse `userRole`, `oneSentenceStory`, and `successCriteria` from its frontmatter and propose them as the primary user and tasks. Ask the user to confirm or extend.

Capture, via `AskQuestion`:
- **Primary user** — specific role and context (e.g. "Maintenance engineers on offshore platforms").
- **2–3 critical tasks** — the workflows this user needs to complete (e.g. "Check pump vibration alerts", "Schedule maintenance work").
- **Context** — experience level, time constraints, device, success criteria.

## Step 2 — Walk each task end-to-end (manual)

Instruct the user to:
1. Open the app **as that user** in a clean browser session with representative test data.
2. Complete each task from beginning to end without shortcuts.
3. Note pain points: where they get stuck, confused, or make errors.

For each task, prompt the user to paste back: what happened, where they got stuck, and any screenshots / notes. Capture these as `taskWalkthroughs[]` for the report.

Do NOT proceed to scoring until the user confirms they walked every task. If they refuse, write a stub report that records "task walkthrough skipped" and exits — do not score.

## Step 3 — Score the 10 questions (probe → propose → confirm)

For every question Q1–Q10, follow the same loop:

1. **Run the listed probes.** They are concrete shell / grep / lint / build commands that produce hard evidence from the repo.
2. **Propose a draft score (1–5)** based on the probe results and the rubric. Show your work: which probe results led to which score.
3. **Cross-check** against the user's task-walkthrough notes from Step 2 (especially for navigation, clickability, error prevention).
4. **Ask the user via `AskQuestion`** with three options: *(a) accept the proposed score `N`*, *(b) override with a specific score*, *(c) override + add a note*.
5. Capture the final score, a one-line rationale, and an improvement note.

### Heuristics for translating probe results into a draft score

These thresholds are starting points — adjust based on the specific evidence and the rubric language. The user always has the final say.

| Signal | Drift toward |
| --- | --- |
| 0 anti-pattern matches, lint clean for the relevant rule | 5 |
| ≤ 3 small matches, mostly in one file | 4 |
| 5–15 matches across several files, or 1 systemic issue | 3 |
| 15+ matches, or pervasive anti-pattern | 2 |
| Anti-pattern is the default style | 1 |

### Per-question automated probes

Each question's probe list is the *first* thing the agent should run before asking the user anything about that question. Always state which probes were run and what they returned.

### The 10 questions and rubric

**Q1 — Aura design system consistency.** Are you using Aura tokens, layouts, components and patterns correctly?

**Probes (automatable):**
- `grep -c '@cognite/aura' package.json` — confirm Aura is a dependency
- `grep -rlE "from '@cognite/aura'" --include='*.ts' --include='*.tsx' src | wc -l` — count files importing Aura
- `grep -rlE '#[0-9a-fA-F]{3,8}' --include='*.css' --include='*.tsx' --include='*.ts' src` — files with hard-coded hex colors
- `grep -rlE '\b(rgb|rgba|hsl|hsla)\(' --include='*.tsx' --include='*.css' src` — files with raw rgb/hsl values
- `npx eslint . --ext .ts,.tsx --rule '{"aura/no-overriding-styles":"error"}' --no-eslintrc --quiet 2>&1 | tail -5` or read the existing lint output for `aura/no-overriding-styles` warning counts

**Translate to draft score:** 0 hard-coded colors + 0 `aura/no-overriding-styles` warnings → 5. Few warnings (1–5) → 4. Many warnings (>15) or no Aura imports → 2–3.

- **5 Excellent:** All Aura tokens applied correctly, no hard-coded values. Proper responsive sizing and page layouts. Aura components used without style overrides. Best practices followed.
- **4 Good:** Mostly Aura tokens and components with 1–2 minor exceptions. Layout spacing mostly consistent. Minimal style overrides.
- **3 Average:** Mix of Aura and custom elements. Some proper spacing, some random values. Overriding styles in multiple places.
- **2 Below average:** Frequently custom colors, typography, or spacing instead of Aura tokens. Heavy customization that breaks patterns.
- **1 Poor:** Not using Aura at all. Custom colors, fonts, spacing throughout.

**Q2 — Navigation, layout and hierarchy.** Can users tell where they are and navigate easily?

**Probes (partially automatable — relies on Step 2 walkthrough):**
- `grep -rcE '(Submit|OK|Click here|Go|Yes|No)]*>[[:space:]]*' --include='*.tsx' src` — empty buttons (icon-only without label needs aria-label, handled in Q10)
- `grep -rlE '` is a smell

**Translate to draft score:** 0 vague labels + every input has a matching label → 5. Few placeholder-only inputs → 4. Vague labels in several places → 3.

- **5:** Every element has a clear, specific label. Plain, action-oriented language ("Save changes", "Delete item").
- **4:** Most labels clear. Minor ambiguity.
- **3:** Labels present but sometimes vague ("Submit", "OK"). Some unnecessary jargon.
- **2:** Many labels unclear. Heavy technical terms without explanation.
- **1:** Labels missing, confusing, or jargon-laden.

**Q4 — System feedback and validation.** Do users know what's happening? Are forms easy to use?

**Probes (automatable):**
- `grep -rlE 'isLoading|isPending|]*onClick' --include='*.tsx' src` — `onClick` on `` (non-semantic, often missing keyboard support)
- `grep -rcE ']*onClick' --include='*.tsx' src` — same for ``
- `grep -rcE 'role="button"' --include='*.tsx' src` — explicit role assignments (good if `` is unavoidable)
- `grep -rcE 'hover:|focus:' --include='*.tsx' src` — Tailwind hover/focus utility usage (high = good)
- `grep -rcE 'cursor-pointer' --include='*.tsx' src` — explicit pointer cursor

**Translate to draft score:** 0 `` without role + many hover/focus utilities → 5. 1–3 violations → 4. Many `onClick` on non-button elements → 2–3.

- **5:** All clickable items look clickable. Hover effects on interactive elements. Cursor changes appropriately.
- **4:** Most interactive elements obvious. Hover effects mostly present.
- **3:** Inconsistent hover states. Occasionally unclear what's interactive.
- **2:** Many interactive elements don't look clickable. Few hover effects.
- **1:** Can't tell what's clickable. No visual feedback.

**Q6 — Error prevention and recovery.** Can users undo or cancel destructive actions?

**Probes (partially automatable):**
- `grep -rilE 'delete|remove|archive|reset' --include='*.tsx' src | head -20` — files with potentially destructive actions
- `grep -rlE 'AlertDialog|ConfirmDialog|window\.confirm' --include='*.tsx' src` — confirm-dialog usage
- `grep -rcE 'variant="destructive"|destructive' --include='*.tsx' src` — destructive button styling
- For each file with destructive verbs, check there is a corresponding `AlertDialog`/`ConfirmDialog` invocation in the same file or its imports

**N/A guidance:** Read-only viewer apps (the common case for Flows demos) have no destructive actions and should score **5 by default with a "no destructive actions" rationale**. Do not penalize an app for not having confirmations it does not need.

- **5:** Confirmation dialogs before destructive actions. Auto-save prevents data loss. Clear undo or cancel options. **OR** the app has no destructive actions.
- **4:** Most destructive actions have warnings. Some auto-save or undo.
- **3:** Some warnings for major actions. Limited undo/cancel.
- **2:** Few warnings. No undo. Easy to lose work.
- **1:** No warnings. No undo. Frequent accidental data loss.

**Q7 — Responsive design and multi-device support.** Does it work on different screen sizes?

**Probes (automatable):**
- `grep -rcE '\b(sm|md|lg|xl|2xl):' --include='*.tsx' src` — Tailwind responsive utility usage (high = good)
- `grep -E ' 0' --include='*.tsx' src` — explicit empty checks

**Translate to draft score:** Every data-fetching panel has an empty-state branch with copy → 5. One or two missing → 4. Many panels missing → 2–3.

- **5:** All empty states show helpful messages and clear next steps. First-time users know exactly what to do.
- **4:** Most empty states helpful. Minor gaps.
- **3:** Some empty states explained. First-time users can figure it out.
- **2:** Many blank pages with no guidance.
- **1:** Blank pages everywhere. No guidance.

**Q9 — Performance and efficiency.** Does the app load quickly?

**Probes (automatable):**

First, check whether a recent build already exists — avoids a slow rebuild when `dist/` is fresh:

```bash
find dist -maxdepth 1 -newer package.json -name '*.js' 2>/dev/null | wc -l
du -sh dist/ 2>/dev/null
```

If the count is 0 (no recent build), fall back to:

```bash
npm run build 2>&1 | tail -20
```

Then gather the remaining metrics:

- `grep -rcE 'React\.lazy|lazy\(' --include='*.tsx' src` — code-split routes (good)
- `grep -rcE 'useMemo|useCallback' --include='*.tsx' src` — memoization usage (informs render efficiency)
- `grep -rlE 'useVirtual|react-window|react-virtual' --include='*.tsx' src` — list virtualization (good for big lists)
- `grep -rlE '\.list\([^)]*\)' --include='*.ts' --include='*.tsx' src | xargs -I{} grep -l 'limit:' {} 2>/dev/null | wc -l` vs total list call sites — pagination coverage
- Cross-reference the latest `code-review-report.md` criterion 2.3 (Limits & pages) score

**Translate to draft score:** Build under 1 MB gzipped + every list has a limit + react-query in use → 5. Bundle 1–2 MB or some lists missing limits → 4. Bundle > 2 MB or systemic unbounded fetches → 2–3.

- **5:** Fast loading with progressive content. Bulk actions, keyboard shortcuts. Common tasks take minimal clicks.
- **4:** Reasonable loading. Most tasks streamlined.
- **3:** Acceptable performance. Tasks moderate effort. Few shortcuts.
- **2:** Slow loading. Tasks require many steps.
- **1:** Very slow or unresponsive.

**Q10 — Accessibility (WCAG AA 2.1).** Can people use it with assistive tech?

**Probes (automatable):**
- Count `` tags and `` tags with `alt` attributes separately to identify missing alt text:
  ```bash
  grep -rcE ']*\balt=' --include='*.tsx' src
  ```
  Any difference means images are missing `alt`.
- `grep -rcE ']*>[[:space:]]*&1 | tail -10`
- If `axe-core` is available: suggest the user run an axe scan in the running app and paste results — automation can flag candidates, not enforce contrast

**Translate to draft score:** 0 missing alts + 0 icon-only buttons without aria-label + focus styles everywhere → 5. A few violations → 4. Systemic gaps → 2–3.

- **5:** All interactions via keyboard. Text contrast meets WCAG AA. Clear focus indicators. Proper ARIA labels. Alt text on images. Touch targets 40px+ / mouse targets 20px+. Form errors announced to screen readers.
- **4:** Most requirements met. Minor exceptions.
- **3:** Basic keyboard support but missing for some features. Mostly acceptable contrast. Focus indicators present but not always clear.
- **2:** Limited keyboard support. Multiple contrast failures. Weak focus indicators.
- **1:** No keyboard navigation. Poor contrast. No focus indicators. Not usable with assistive tech.

## Step 4 — Compute average and quality level

Average = sum of all 10 scores ÷ 10.

Map to the quality level table from the docs:

| Average | Quality level | Recommendation |
| --- | --- | --- |
| 4.5 – 5.0 | Excellent — ready to launch | Minor improvements over time |
| 3.8 – 4.4 | Good — launch with minor fixes | Address lower-scoring areas |
| 3.0 – 3.7 | Average — needs improvement | Fix major problems before launching |
| Below 3.0 | Needs significant work | Substantial improvements required |

`flows-external-app-submit` gates on **average ≥ 3.8**.

## Step 5 — Write the report

Create `reviews/design-review/feedback-round-/design-review-report.md` with this structure:

```markdown
# Design Review —  — round 

## User and tasks

- **Primary user:** ...
- **Tasks evaluated:**
  1. ...
  2. ...
  3. ...
- **Context:** ...

## Task walkthrough findings

- **Task 1 — ...** ...
- **Task 2 — ...** ...
- **Task 3 — ...** ...

## Scores

| Question | Score | Rationale | Improvement note |
| --- | --- | --- | --- |
| Q1 Aura consistency | n | ... | ... |
| Q2 Navigation & hierarchy | n | ... | ... |
| Q3 Labels & language | n | ... | ... |
| Q4 Feedback & validation | n | ... | ... |
| Q5 Clickability | n | ... | ... |
| Q6 Error prevention | n | ... | ... |
| Q7 Responsive | n | ... | ... |
| Q8 Empty states | n | ... | ... |
| Q9 Performance | n | ... | ... |
| Q10 Accessibility | n | ... | ... |

## Summary

- Average score: 
- Quality level: 

## Must Fix (any score < 3)

- ...

## Should Fix (any score 3 – 3.7)

- ...

## Nice to Fix (any score 3.8 – 4.4)

- ...
```

The `Average score:` line must be machine-readable in exactly that format — `flows-external-app-submit` parses it.

## Step 6 — Print the gate status

After writing, print to the terminal:
- The average score
- The quality level
- Whether the result meets the `flows-external-app-submit` gate (≥ 3.8)
- If below 3.8, instruct the user to fix Must Fix and Should Fix items and re-run this skill in a new feedback round.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [cognitedata](https://github.com/cognitedata)
- **Source:** [cognitedata/builder-skills](https://github.com/cognitedata/builder-skills)
- **License:** Apache-2.0

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-cognitedata-builder-skills-flows-design-review
- Seller: https://agentstack.voostack.com/s/cognitedata
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
