# Data Architecture Design

> Create data architecture design documents for a specific canonical data object, including data flow diagrams, integration traceability, ownership, lifecycle, quality, privacy, and governance. Use when the user asks to design, document, review, or update the architecture of a data object and link it to Target Architecture Phase C.

- **Type:** Skill
- **Install:** `agentstack add skill-future-cx-ai-architecture-toolkit-data-architecture-design`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [Future-CX](https://agentstack.voostack.com/s/future-cx)
- **Installs:** 0
- **Category:** [Agent Skills](https://agentstack.voostack.com/c/agent-skills)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [Future-CX](https://github.com/Future-CX)
- **Source:** https://github.com/Future-CX/AI-Architecture-Toolkit/tree/main/skills/data-architecture-design

## Install

```sh
agentstack add skill-future-cx-ai-architecture-toolkit-data-architecture-design
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# Data Architecture Design

## Quick Start

Assume the Data Architect Agent role from `agents/data-architect.md`: clarify the canonical data object, ownership, source of truth, lifecycle, quality, privacy, retention, integration flows, and governance expectations.

Before creating or updating any data architecture design files, run a `grill-me` clarification session using `../grill-me/SKILL.md`. Ask one question at a time until the target architecture link, data object meaning, ownership, lifecycle, integrations, quality expectations, privacy concerns, assumptions, and open questions are clear enough to avoid avoidable misunderstanding.

During that clarification session, use `../ubiquitous-language/SKILL.md` whenever terms are vague, missing, overloaded, conflicting, or important enough to become shared domain language. Create or update `/GLOSSARY.md` inline as terms are clarified; do not batch glossary updates until the end.

Use `templates/data-architecture-design-template.md` as the output structure. Replace placeholders and drafting guidance with concrete content; mark unknown facts as `TBD` or open questions.

Store generated data architecture designs under the consuming repository's private lab root:

```text
data-architectures//-data-architecture-design.md
data-architectures//diagrams/
```

Do not write real-company data architecture details into this public toolkit repository.

## Required Inputs

- Target architecture document to link to
- Canonical data object name
- Data object description and business purpose
- Main capability or business process using the data object
- Source of truth and data owner
- Data classification
- Producing systems, consuming systems, and integration touchpoints
- Data lifecycle states, retention, privacy, and classification expectations
- Known quality rules, reconciliation needs, lineage needs, and governance constraints

## Workflow

1. Ask which target architecture document this data architecture design supports. Read it when a path is provided.
2. Start a mandatory `grill-me` clarification session using `../grill-me/SKILL.md`. Ask one question at a time and cover:
   - Canonical data object meaning and aliases
   - Business purpose and main capability
   - Source of truth, data owner, steward, allowed writers, and consumers
   - Producing systems, consuming systems, integrations, and transformations
   - Lifecycle states, retention, archival, deletion, audit, and exception handling
   - Quality rules, reconciliation, lineage, observability, and operational ownership
   - Privacy, classification, access, residency, masking, encryption, and compliance constraints
   - Assumptions, unresolved decisions, and open questions
3. During the `grill-me` session, validate terminology against `/GLOSSARY.md`. If the data object, applications, capabilities, integrations, lifecycle states, or important terms are missing or ambiguous, use `../ubiquitous-language/SKILL.md` immediately to update the private lab glossary.
4. After the clarification session, validate that `/GLOSSARY.md` was created or updated during the current run. If it was not created or updated, stop before generating data architecture files and ask the user to run `grill-me` followed by `ubiquitous-language` so the glossary is updated first.
5. Ask for the canonical data object, source of truth, owner, main capability, producers, consumers, and known integrations when still not provided.
6. Determine the output folder as `/data-architectures//`.
7. Create or update `-data-architecture-design.md` from `templates/data-architecture-design-template.md`. Preserve the two opening tables: the document metadata table first, followed by the `Data Architecture Overview` table.
8. Create diagrams that make the data movement understandable:
   - Data flow diagram from `../create-drawio-diagram/templates/data-flow.drawio`
   - Data Architecture Design diagram from `../create-drawio-diagram/templates/data-architecture-diagram.drawio`, showing systems, interfaces, events, files, APIs, or batches
   - Optional conceptual data model when the object has important relationships
9. Create every Draw.io diagram with `../create-drawio-diagram/SKILL.md` because it contains the required diagram instructions, style rules, and templates. Store each editable `.drawio` source in `diagrams/`.
10. Export a same-basename `.svg` file for each Draw.io diagram that must be embedded, embed the SVG in the document, and link the `.drawio` source near the embedded SVG.
11. Link the data architecture design from Phase C of the target architecture:
   - Add it to the `## Data Architecture Designs` table in `05-phase-c-data-architecture.md` when section files exist.
   - Also update the assembled `target-architecture-document.md` when it exists.
12. Populate `## Relevant Links` with the target architecture, capability overview, solution architecture design, integration designs, ADRs, and glossary references that are explicitly related. Use linked document titles as Markdown link labels when available.
13. Run the glossary and readability gate before final delivery:
   - Read `/GLOSSARY.md`.
   - In the Glossary, find the `Jargon` section and its Avoid list.
   - Replace avoidable jargon in stakeholder-facing sections: `## Short Summary`, `## Description`, `## Data Flow`, `## Ownership and Source of Truth`, `## Lifecycle`, `## Data Quality and Lineage`, and `## Privacy, Security, and Compliance`.
   - Run the `check-readability` skill in `../check-readability/SKILL.md` on the data architecture design.
   - Update the top-table `Readability Score` row with the rounded numeric Flesch Reading Ease score.
   - Aim to improve the full document toward a Flesch Reading Ease score of 40-50, while allowing lower scores when architecture documents require specialist terms.
   - If the score remains below 40, make sure there are no sentences over 30 words in stakeholder-facing sections and explain any unavoidable specialist terms.
14. Capture unresolved ownership, lineage, quality, retention, privacy, integration, and operational facts as open questions rather than inventing details.

## Writing Guidance

- Center the document on one canonical data object.
- Start with the document metadata table, then the `Data Architecture Overview` table with the data object, source of truth, main capability, data owner, and classification.
- Include a `Short Summary` section before `Description` that gives the business meaning, ownership, and main architecture concern in a few plain-language sentences.
- Write stakeholder-facing sections in plain business language first: `Short Summary`, `Description`, `Data Flow`, `Ownership and Source of Truth`, `Lifecycle`, `Data Quality and Lineage`, and `Privacy, Security, and Compliance`.
- Prefer short sentences and active voice. Use technical terms only when they affect ownership, risk, integration behavior, support, privacy, or a decision.
- Keep dense implementation detail in tables or linked integration designs.
- Treat the data architecture design as the source of truth for the canonical data model. Solution architecture documents should reference this model and describe application-specific usage, not redefine it.
- Be explicit about source of truth, authoritative owner, allowed writers, and downstream consumers.
- Identify data lifecycle states, create/update/delete behavior, retention, archival, purge, and audit requirements.
- Describe data movement in business terms first, then technical integration details.
- Link related integration designs instead of duplicating full interface contracts.
- Document data quality rules, validation points, reconciliation, lineage, observability, and stewardship responsibilities.
- Include privacy and security concerns such as classification, sensitive attributes, access controls, consent, residency, masking, encryption, and audit logging.
- Keep relevant links bidirectional. If the data architecture design is linked from Target Architecture Phase C, link back to the target architecture from the data architecture design.
- Use linked document names as link labels. For local Markdown files, derive the name from the first `#` heading; otherwise use the filename without extension.
- Maintain `/GLOSSARY.md` while writing whenever the design introduces or changes domain terms, applications, canonical data objects, lifecycle states, integration names, ownership roles, relationships, jargon, deprecated terms, or words to avoid.

## Readability Requirements

Data architecture designs are stakeholder-facing architecture documents. Architects and delivery teams need enough detail to act, but the opening and ownership sections must be understandable to non-technical business stakeholders.

Use the `check-readability` skill in `../check-readability/SKILL.md` before finishing a data architecture design. Treat the target audience as non-technical business stakeholders unless the user gives a different audience. Aim to improve the Flesch Reading Ease score toward 40-50, but allow lower scores when required architecture terms, legal terms, data object names, system names, or quoted source text make that target impractical.

The data architecture design is not complete until the readability check confirms these expectations:

- The top metadata table's `Readability Score` row contains the rounded numeric Flesch Reading Ease score.
- Glossary `Jargon` terms from `/GLOSSARY.md` are removed from stakeholder-facing sections unless they are quoted source text or explicit glossary references.
- Preferred glossary terms are used exactly when the glossary gives one. If a preferred term is missing or unclear, use plain language and capture the terminology gap as an open question.
- Acronyms and specialist terms are explained the first time they appear unless they are already defined in `GLOSSARY.md`.
- If the score remains below 40, stakeholder-facing sections have no sentences over 30 words and any unavoidable specialist terms are briefly explained.
- Long table cells are moved into notes, linked integration designs, or later technical sections when they make ownership, quality, lifecycle, privacy, risks, or decisions hard to scan.

## Data Flow Diagram Format

Use `../create-drawio-diagram/templates/data-flow.drawio` as the starting point for the data flow diagram. The diagram should read like an operational trace of the data object across systems and process steps.

- Title the diagram ` | Data Flow | `.
- Put the business journey, process stages, screens, or major events across the top from left to right when they are known.
- Use one horizontal lane per concrete solution or component. Keep Customer first, Channel second, optionally add one backend-for-frontend lane when a BFF participates in the flow, then Engagement solution lanes, Integration component lanes, and Enterprise Foundation or MDM solution lanes.
- Do not group several solutions into one broad lane. Add more vertical canvas space instead.
- Draw data movements as vertical or orthogonal arrows crossing lanes. Label each arrow with the specific data object, event, command, file, API call, batch, or transformation.
- Use color intentionally:
  - Blue for primary read, write, replication, or publication flows.
  - Green for enrichment, rules, calculation, validation, or decisioning flows.
  - Grey or dashed lines for optional, planned, deprecated, or uncertain flows.
- Show where the data object is created, updated, enriched, read, replicated, archived, deleted, or submitted.
- Keep lane labels stable and readable on the left. Keep each process-stage label centered over the boxes and arrows that belong to that stage.
- Prefer a wide landscape canvas over compressed diagrams. Increase the canvas width for more stages and the canvas height for more solution lanes until arrows, labels, lane headers, stage labels, and boxes do not overlap.
- Do not use a generic box-and-line context view for the data flow. The data flow diagram must show movement through lanes over time or process progression.

## Data Architecture Design Diagram Format

Use `../create-drawio-diagram/templates/data-architecture-diagram.drawio` as the starting point for the Data Architecture Design diagram. Follow the `Data Architecture Diagram Layout` rules in `../create-drawio-diagram/SKILL.md`.

Adapt the data architecture overview elements to show the data object's source systems, canonical data object, owners, consumers, integrations, governance touchpoints, and external systems. Always place Backend-for-Frontend components in the Frontend layer. Place different peer components horizontally instead of stacking them vertically unless they are part of the same direct end-to-end flow. Keep the diagram focused on integration traceability for the data object; put interface detail in the Data Architecture Design table or linked integration designs.

## Phase C Link Format

Use this table shape in Target Architecture Phase C:

```md
## Data Architecture Designs

| Data Architecture Design                                                | Data Object     | Source of Truth     | Description     |
| ----------------------------------------------------------------------- | --------------- | ------------------- | --------------- |
| [{{DATA_ARCHITECTURE_DESIGN_TITLE}}]({{DATA_ARCHITECTURE_DESIGN_LINK}}) | {{DATA_OBJECT}} | {{SOURCE_OF_TRUTH}} | {{DESCRIPTION}} |
```

## Guardrails

- Keep real-company data architecture designs in a private company lab repository.
- Write generated files to `/data-architectures/`, not inside this public toolkit repository.
- Do not create or update a data architecture design without an explicit target architecture link.
- Do not create or update a data architecture design without first running the `grill-me` clarification session and updating the private lab glossary with `ubiquitous-language` when terminology is missing or unclear.
- Do not invent internal system names, data fields, classifications, retention periods, integration contracts, or non-public business context.
- Do not overwrite existing documents or diagrams unless the user explicitly asks.

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [Future-CX](https://github.com/Future-CX)
- **Source:** [Future-CX/AI-Architecture-Toolkit](https://github.com/Future-CX/AI-Architecture-Toolkit)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-future-cx-ai-architecture-toolkit-data-architecture-design
- Seller: https://agentstack.voostack.com/s/future-cx
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
