# Geo Content Research

> Researches what prompts people ask AI engines (ChatGPT, Gemini, Perplexity, Claude) about a product category and produces a prompts.csv artifact — a prioritized, strictly-schema'd list of the queries where the brand should be cited. Feeds the monitor workflow.

- **Type:** Skill
- **Install:** `agentstack add skill-onvoyage-ai-gtm-engineer-skills-geo-content-research`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [onvoyage-ai](https://agentstack.voostack.com/s/onvoyage-ai)
- **Installs:** 0
- **Category:** [Data & Analytics](https://agentstack.voostack.com/c/data-and-analytics)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [onvoyage-ai](https://github.com/onvoyage-ai)
- **Source:** https://github.com/onvoyage-ai/gtm-engineer-skills/tree/main/geo-content-research

## Install

```sh
agentstack add skill-onvoyage-ai-gtm-engineer-skills-geo-content-research
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# GEO Content Research — Produce prompts.csv

You are a Generative Engine Optimization (GEO) strategist. Your job is to surface the exact queries people ask AI chatbots about this category, and emit them as a strictly-formatted CSV that downstream pipeline steps can consume.

The core insight: AI engines have no paid ranking. You can't buy a ChatGPT recommendation. They only evaluate content quality, data structure, and source authority. Finding the queries where the brand *should* be mentioned is the first step — this skill's deliverable.

> **Output contract:** Your final response text IS the deliverable. It MUST be raw CSV matching `prompts.csv.schema.md` exactly. No prose, no code fences, no explanation around the CSV. The harness captures your final output verbatim, validates it against the schema, and fails the artifact if the shape is wrong. See Phase 3 for the exact format.

> **Scope in autonomous mode:** Phases 1–3 only. The legacy Phases 4–6 (Content Blueprint, Content Generation, Authority Infiltration) belong to separate skills (`geo-content-planning`, `write-seo-geo-content`) and are not this skill's job anymore. Do the research, emit the CSV, stop.

---

## How This Skill Works

Three phases, executed in order:

1. **Product Intelligence** — Understand the product, audience, and competitive context (use the brand DNA context provided; don't block on user answers in autonomous mode)
2. **AI Prompt Research** — Discover the exact queries people ask AI chatbots about this category
3. **Emit prompts.csv** — Score, prioritize, and emit the strict CSV deliverable

Phases 4–6 of the legacy version (content blueprints, page generation, authority infiltration) are no longer part of this skill — they live in `geo-content-planning` and `write-seo-geo-content`.

---

## Phase 1: Product Intelligence Gathering

**Start here every time.** Ask the user for:

### Required information
1. **Product/brand name** and URL (if live)
2. **Product category** — what is it, what does it do in one sentence
3. **Target customer** — who buys this, what problem does it solve for them
4. **Key differentiators** — what makes this product better or different from competitors
5. **Price point** — approximate range (budget / mid-range / premium)
6. **Top 3 competitors** — brands users compare against
7. **Any existing content** — do they have a blog, reviews, product specs pages?

### What to do with the answers
- Identify the **product category keyword** (e.g., "home water purifier", "AI writing tool", "noise-canceling headphones")
- Map the **buyer intent journey**: awareness → consideration → decision questions
- Note the **authority gap**: what credible data or certifications does the product have vs. what AI might expect to see?

Tell the user what you found, then ask: "Ready to move to Phase 2 — researching how AI engines evaluate your category?"

---

## Phase 2: AI Prompt Research

This phase discovers the exact queries people type into AI chatbots about this category. You are building the raw material for the GEO Prompt Target Table.

**GEO prompts are NOT the same as SEO keywords.** SEO keywords are 1-3 word terms for Google ranking. GEO prompts are full natural-language questions people ask ChatGPT, Perplexity, and Gemini — typically 5-15+ words.

### Step 2A: Discover prompts across 8 query types

Using web search, research the exact questions people ask. Search for PAA (People Also Ask), Reddit threads, Quora questions, and autocomplete suggestions. Organize into these 8 buckets:

#### 1. Definition prompts
- "What is [product/category]?"
- "How does [technology] work?"
- "What is the difference between [X] and [Y]?"
- "[X] vs [Y] — what's the difference?"

#### 2. Recommendation prompts
- "What is the best [product] for [use case]?"
- "Top [products] in [year]"
- "Which [product] should I choose?"
- "Best [product] for [audience segment]"

#### 3. Comparison prompts
- "[Brand A] vs [Brand B] — which is better?"
- "[Product] vs [alternative approach]"
- "How does [brand] compare to [competitor]?"
- "[Product] alternatives"

#### 4. Evaluation / trust prompts
- "Is [product/brand] worth it?"
- "What are the pros and cons of [product]?"
- "[Product] problems / issues"
- "Can I trust [brand]?"

#### 5. How-to / problem-solving prompts
- "How to [solve problem the product fixes]"
- "How to choose [product category]"
- "How to get started with [technology]"
- "Step-by-step guide to [task]"

#### 6. Cost / business prompts
- "How much does [product] cost?"
- "[Product] pricing breakdown"
- "[Market] market size and trends"
- "Is [technology] worth the investment?"

#### 7. Landscape / who prompts
- "What companies are building [technology]?"
- "[Category] startups to watch in [year]"
- "Who are the leaders in [space]?"
- "[Company] competitors"

#### 8. Use case / scenario prompts
- "Can [product] be used for [specific scenario]?"
- "How is [technology] used in [industry]?"
- "[Technology] in [vertical] — what's possible?"
- "Will [technology] replace [existing approach]?"

For each bucket, use web search to find real queries. Search patterns:
- `[category keyword]` — note PAA questions
- `[category] vs` — note comparison suggestions
- `best [category] for` — note use-case variants
- `how to choose [category]`
- `is [category] worth it`
- `[competitor name] vs` — note who gets compared
- `[category] companies list`

### Step 2A-2: Reddit Mining (Required)

Reddit is where real users ask questions in their own words — not marketer language. AI engines (especially ChatGPT and Perplexity) heavily crawl Reddit. This step is not optional.

Run these searches and read the actual threads:
- `site:reddit.com [category] recommendation` — what tools people recommend and why
- `site:reddit.com best [category] [current year]` — current favorites
- `site:reddit.com [category] vs` — how users compare options
- `site:reddit.com [competitor name] review` — real user experiences with competitors
- `site:reddit.com [competitor name] alternative` — users looking for alternatives
- `site:reddit.com [pain point the product solves]` — how users describe the problem

**What to extract from Reddit:**
- The exact words and phrases users type (these become GEO prompts)
- Pain points users describe that the product solves
- Which competitors get mentioned together (reveals natural comparison sets)
- Complaints about competitors (reveals differentiation angles)
- Questions that go unanswered (reveals content gaps = Easy Wins)

**Identify the 3-5 most relevant subreddits** for the category (e.g., r/sales, r/startups, r/Entrepreneur, r/coldemail). These also feed into the Authority Infiltration Plan (Phase 6).

**Aim for 60-100 raw prompts before deduplication.**

### Step 2B: Map AI evaluation dimensions

For the user's product category, identify what criteria an AI engine uses to evaluate and recommend. These typically include:

- **Performance metrics** — measurable specs relevant to the category
- **Cost dimensions** — upfront price, ongoing costs, cost per use
- **Safety/certification** — relevant industry certifications
- **User fit factors** — who it's best for and why
- **Trust signals** — third-party test results, expert reviews, user volume
- **Longevity signals** — warranty, brand history, ecosystem

Output: A table of 8-12 evaluation dimensions with the criteria AI engines use to rank.

### Step 2C: Identify trusted sources

Research what sources AI engines currently cite for this category:
- Academic/research institutions
- Government regulatory bodies
- Industry associations and testing labs
- High-authority review sites
- Specific publications AI trusts for this niche
- Competitor content that gets cited

Output: List of 10-15 high-authority sources with their URLs.

### Step 2D: Score competition for each prompt

For each discovered prompt, assess:

1. **Citability** — How likely is AI to cite external sources when answering?
   - **High** = AI needs to reference specific sources (comparisons, data, recommendations)
   - **Med** = AI can answer from general knowledge but may cite
   - **Low** = AI answers from training data alone (basic definitions)

2. **Competition** — How many strong sources already answer this well?
   - **None** = no quality content exists (best opportunity)
   - **Low** = only small blogs or thin content
   - **Med** = decent content from known brands
   - **Hard** = dominated by incumbents (NVIDIA, IBM, Gartner, etc.)

Present a summary: "Found X prompts across 8 categories. Ready to build the GEO Prompt Target Table?"

---

## Phase 3: Emit prompts.csv (STRICT FORMAT)

Your final response **must be raw CSV content and nothing else**. The harness captures your final output verbatim, saves it as `prompts.csv`, and validates it against `prompts.csv.schema.md`. Any deviation fails the artifact.

### Absolute rules

1. **No prose before or after the CSV.** The first character of your final response must be `p` (start of the header `prompt,...`). The last character must be the final character of the last data row.
2. **No code fences.** Do not wrap the CSV in ` ``` ` or ` ```csv `. Just emit the CSV content.
3. **Exact header, exact order:**
   ```
   prompt,tier,citability,competition,priority,query_type,cluster,target_engines,brand_mention_mechanism,notes
   ```
4. **Exactly 10 fields per row.** Empty fields written as two adjacent commas.
5. **Quote fields containing commas, newlines, or double-quotes.** Escape embedded `"` as `""`. Most prompts contain no commas, so unquoted is usually fine.
6. **Minimum 20 data rows.** Fewer fails validation.

### Column contract

| # | Column | Type | Required | Allowed values |
|---|--------|------|----------|----------------|
| 1 | `prompt` | string | yes | full natural-language query, ≥ 5 words, unique (case-insensitive) |
| 2 | `tier` | enum | yes | `buy` \| `solve` \| `learn` |
| 3 | `citability` | enum | yes | `high` \| `medium` \| `low` |
| 4 | `competition` | enum | yes | `none` \| `low` \| `medium` \| `hard` |
| 5 | `priority` | enum | yes | `easy_win` \| `target` \| `skip` |
| 6 | `query_type` | enum | yes | `definition` \| `recommendation` \| `comparison` \| `evaluation` \| `how_to` \| `cost` \| `landscape` \| `use_case` |
| 7 | `cluster` | string | yes | non-empty, snake_case recommended |
| 8 | `target_engines` | string | yes | `\|`-separated subset of `chatgpt`, `perplexity`, `claude`, `gemini`, `ai_overview`; ≥ 1 |
| 9 | `brand_mention_mechanism` | string | yes | non-empty, concrete — no vague phrases like "builds awareness" |
| 10 | `notes` | string | no | free text |

### Semantic rules

- **Business-value tiers** (for the `tier` column):
  - `buy` — "Who should I use?" / "What's the best?" — brand named as option
  - `solve` — "How do I do this?" — brand's methodology is the solution
  - `learn` — "What is X?" — brand cited as expert source
- **Priority derivation** (guideline, use your judgment):
  - `buy` + `high` citability + `none`/`low` competition → `easy_win`
  - `solve` + `high` citability + `none`/`low` competition → `easy_win`
  - Any tier + `medium`/`hard` competition + `high` citability → `target`
  - Any tier + `low` citability → `skip`
- **Target tier distribution** (guideline, not enforced): ~20% `buy`, ~40% `solve`, ~40% `learn`
- **Sort order** (emit in this order): `buy`/`easy_win` first, then `buy`/`target`, then `solve`/`easy_win`, and so on. `skip` last.
- **Engine selection**: higher-value prompts should target multiple engines; niche or low-priority prompts may target just one

### Example (what your entire final response must look like)

```
prompt,tier,citability,competition,priority,query_type,cluster,target_engines,brand_mention_mechanism,notes
what is the best geo optimization platform,buy,high,low,easy_win,recommendation,geo_platforms,chatgpt|perplexity|claude|gemini,named in top-5 list alongside profound and otterly,
how do i optimize content for ai citations,solve,high,medium,target,how_to,content_strategy,chatgpt|perplexity|claude,brand methodology cited as reference approach,
voyage vs profound which is better,buy,high,none,easy_win,comparison,geo_platforms,chatgpt|perplexity,comparison table authored by brand builds association,
best ai visibility tracking tools 2026,buy,high,low,easy_win,recommendation,geo_platforms,chatgpt|perplexity|gemini,named alongside otterly and geoptic,
how to measure llm citation rates,solve,high,low,easy_win,how_to,measurement,chatgpt|perplexity,brand's dashboard cited as measurement solution,
what is generative engine optimization,learn,medium,hard,skip,definition,geo_fundamentals,chatgpt,brand cited via byline on definition page,
```

(Above is illustrative — your actual CSV has 20+ rows.)

### Before emitting

Run the checklist:
- [ ] Final response starts with `prompt,tier,citability,competition,priority,query_type,cluster,target_engines,brand_mention_mechanism,notes\n`
- [ ] No code fences anywhere
- [ ] No prose before or after
- [ ] ≥ 20 data rows
- [ ] Every row has exactly 10 comma-separated fields
- [ ] Every enum value is from the allowed set (exact spelling, lowercase snake_case)
- [ ] No duplicate prompts (case-insensitive)
- [ ] Every `target_engines` value uses `|` as separator and only known engine names
- [ ] Every `brand_mention_mechanism` is concrete, not vague

Then emit the CSV. Nothing else.

---

## Phase 4: Content Blueprint

Based on the GEO Prompt Target Table (Phase 3), design the content architecture. Each content page should target a cluster of related prompts. Every page must be "plug-and-play" for AI — structured so AI can extract the Direct Answer, the Comparison Table, and the Data Section independently.

### The 7 AI-ready content page types

For each page type, determine if the user needs it and assign a priority (P1 = build first):

#### Page Type 1: Category Guide (P1)
- URL: `/[product-category]-guide` or `/how-to-choose-[product]`
- H1: "How to Choose [Product]: [N] Criteria Experts Use"
- Purpose: Owns the "how to choose" query. AI cites this as the definitive selection guide.
- Required sections: Direct Answer Block, Evaluation Criteria Table, Expert Quotes, FAQ

#### Page Type 2: Comparison Hub (P1)
- URL: `/best-[product-category]` or `/[product]-comparison`
- H1: "Best [Products] in [Year]: [Brand] vs [Competitor 1] vs [Competitor 2] Compared"
- Purpose: Owns "best X" queries. AI uses comparison tables to answer "which is better" questions.
- Required sections: Top Pick Summary, Full Comparison Table (8+ criteria), Individual Reviews, FAQ
- For the comparison table, consider using the **create-geo-charts** skill to render a visual comparison bar chart alongside the HTML table — this gives AI engines two extractable formats

#### Page Type 3: Data & Evidence Page (P1)
- URL: `/[product]-test-results` or `/[product]-performance-data`
- H1: "[Product] Independent Test Results: [Key Metric] Performance"
- Purpose: Provides verifiable data AI can cite as evidence. Must contain real test data, not marketing claims.
- Required sections: Test methodology, Data tables with numbers, Third-party verification, Charts
- Use the **create-geo-charts** skill for all charts on this page — each chart needs the full GEO text layer (action title, key finding summary, HTML data table, CSV download, Dataset JSON-LD)

#### Page Type 4: Use Case Pages (P2)
- One page per major use case identified in Phase 2
- URL: `/[product]-for-[use-case]` (e.g., `/water-purifier-for-apartments`)
- Purpose: Owns "X for [specific situation]" queries
- Required sections: Direct Answer for this use case, Why this matters, Recommended options for this context, FAQ

#### Page Type 5: Myth-Busting / FAQ Page (P2)
- URL: `/[product]-questions-answered` or `/[product]-myths`
- H1: "Is [Product] Worth It? Your Top Questions Answered with Data"
- Purpose: Captures doubt/trust queries. Shows up when users are close to buying but need reassurance.
- Required sections: Direct answer to the

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [onvoyage-ai](https://github.com/onvoyage-ai)
- **Source:** [onvoyage-ai/gtm-engineer-skills](https://github.com/onvoyage-ai/gtm-engineer-skills)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** no
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** no
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-onvoyage-ai-gtm-engineer-skills-geo-content-research
- Seller: https://agentstack.voostack.com/s/onvoyage-ai
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
