# Open Graph Share Seo Geo

> How to make a page unfurl reliably in iMessage, WhatsApp, Slack, Discord, LinkedIn, and X; surface to search engines via sitemap.xml + robots.txt; and stay legible to generative engines (GEO), including the llms.txt standard for LLM corpus ingest. Use when adding or debugging OpenGraph / Twitter Card metadata, picking an OG image format, choosing where to host the image, fixing pages that "won't…

- **Type:** Skill
- **Install:** `agentstack add skill-lossless-group-lossless-agent-skills-open-graph-share-seo-geo`
- **Verified:** Yes — security-reviewed for prompt injection and unsafe behavior
- **Seller:** [lossless-group](https://agentstack.voostack.com/s/lossless-group)
- **Installs:** 0
- **Category:** [AI & ML](https://agentstack.voostack.com/c/ai-and-ml)
- **Latest version:** 0.1.0
- **License:** MIT
- **Upstream author:** [lossless-group](https://github.com/lossless-group)
- **Source:** https://github.com/lossless-group/lossless-agent-skills/tree/main/open-graph-share-seo-geo

## Install

```sh
agentstack add skill-lossless-group-lossless-agent-skills-open-graph-share-seo-geo
```

Requires the [AgentStack CLI](https://agentstack.voostack.com/docs/cli). Works with Claude Code, Cursor, and any MCP-compatible agent.

## About

# OpenGraph, Share, SEO & GEO

> Make the link unfurl. Make it fast. Make it match the bytes.

Lossless Group convention for share metadata across all sites (Astro Knots, plugin pages, splash, fundraise decks). Optimized for **iMessage and WhatsApp first** (the channels we share through most), with Slack, Discord, LinkedIn, X, Facebook, and generative engines (Perplexity, ChatGPT, Claude, Gemini) as concentric rings around that.

## When to use this skill

- Adding or auditing `` meta tags on any page that gets shared
- "It's not unfurling in iMessage / WhatsApp / Slack" debugging
- Picking the OG image format (`.webp` vs `.jpg` vs `.png`)
- Deciding where to host the OG image (`/public` vs CDN)
- Building or modifying a `MetaTags`-style component
- Reviewing or shipping a marketing splash, plugin page, blog post, fundraise deck
- Setting up GEO (Generative Engine Optimization) on a content-heavy page
- Scaffolding `/llms.txt` and `/llms-full.txt` for LLM corpus ingest, or porting that pattern to a new splash
- Adding `@astrojs/sitemap` + `robots.txt` to a splash or marketing site for proper search-engine discoverability

## Hard rules (Why + How to apply)

### 1. Host the OG image on a remote CDN, not `/public`

**Rule:** Default OG image goes on a CDN (ImageKit, Cloudflare Images, S3+CloudFront, Vercel Blob). Do not point `og:image` at a local `/public/og.png` served by GitHub Pages or a static host.

**Why:** Local assets behind GitHub Pages unfurl intermittently in iMessage, WhatsApp, and Slack. The original page renders fine, but the unfurler silently skips the image — usually because of slow first-byte time, missing `Content-Length`, missing CORS, or aggressive negative caching. CDNs (ImageKit in particular) ship the right headers (`access-control-allow-origin: *`, `cache-control: public, max-age=…`, accurate `content-length`, `etag`) and serve from edges close to the unfurler.

**How to apply:** When scaffolding a new site's OG defaults, upload the banner to ImageKit first and reference the absolute URL. Treat `/public/og-default.png` as a fallback only — better yet, do not bother committing it.

### 2. Prefer JPEG over WebP for the bytes the unfurler receives

**Rule:** The image bytes that reach iMessage, WhatsApp, Slack, Discord, LinkedIn, X, and Facebook should be **JPEG**. PNG is a fine second choice for graphics with text. WebP is risky.

**Why:** WebP support across unfurlers is uneven and historically silent-failing. iMessage and WhatsApp have both shipped versions that ignore WebP previews entirely. JPEG is universally accepted, ~95 KB at 1200×630 for a photographic banner — small enough that no unfurler chokes on it.

**How to apply:** If you are exporting from Image-Gin / Figma / Photoshop, export JPEG. If you are using ImageKit transformations, request `?tr=f-jpg` or rely on its content negotiation (see rule 3). To convert an existing local file, the user is likely to have `ffmpeg` installed — `ffmpeg -i in.webp out.jpg` is the simplest one-liner; see [references/seo-best-practices.md](references/seo-best-practices.md) for more conversion recipes.

### 3. `og:image:type` must match the actual bytes the unfurler receives — not the URL extension

**Rule:** Verify with `curl -sI` (no `Accept` header) what `content-type` the server returns. Whatever that is — `image/jpeg`, `image/png`, `image/webp` — is what goes in ``. The URL's file extension is not authoritative.

**Why:** ImageKit (and many modern CDNs) content-negotiate via `Vary: Accept`. A URL ending in `.webp` will serve `image/webp` to a Chrome browser that sends `Accept: image/webp`, but `image/jpeg` to an unfurler that doesn't. If your meta tag declares `image/webp` but the unfurler downloads `image/jpeg`, strict validators (and some unfurlers) bail. Match the declared type to the bytes most unfurlers actually receive.

**How to apply:** Before committing, run:
```bash
curl -sI "" | grep -iE "content-type|content-length|vary"
```
The `content-type` line is what you put in `og:image:type`. If `vary: accept` is present, you are content-negotiated — declare the JPEG/PNG response (the no-`Accept` default), not the WebP variant.

### 4. Absolute URLs only

**Rule:** Every `og:image`, `og:image:secure_url`, `og:url`, `twitter:image`, and canonical `href` is a fully qualified `https://…` URL.

**Why:** Unfurlers do not know the origin. A path like `/og.png` resolves to *their* origin (often nothing) and is dropped. This bites especially hard on sites served under a base path (GitHub Pages `/repo-name/`).

**How to apply:** In Astro, build the absolute URL from `Astro.site` (or `Astro.url.origin`) plus `import.meta.env.BASE_URL` plus the path. Branch on `startsWith('http')` so absolute URLs pass through untouched. See the `MetaTags.astro` pattern in any Astro Knots site.

### 5. Always emit the full image-meta sextet

**Rule:** Every page emits all six OG image properties:

```html

```

Plus the Twitter card pair:
```html

```

**Why:** Width and height let the unfurler reserve layout space without downloading the image. Type lets it skip formats it cannot render. Alt is required for accessibility *and* used by some clients (Slack) as the fallback caption. `secure_url` is legacy but still consulted by older clients. Omitting any of these turns into "sometimes it shows, sometimes it doesn't."

### 6. Standard banner dimensions: 1200 × 630

**Rule:** Default OG banners are **1200 × 630** (the "Image-Gin wide" standard). Plugin/page-specific overrides may use other ratios but must declare matching width/height in the meta tags.

**Why:** 1200 × 630 is the largest size Facebook/LinkedIn/X render without re-cropping, the size iMessage and WhatsApp expect for "rich" previews, and small enough (~90–150 KB JPEG) for fast unfurling. 1408 × 704 and similar non-standard sizes get center-cropped or downscaled inconsistently.

**How to apply:** Image-Gin is configured to export 1200 × 630. If you are hand-cropping in Figma, snap to that size. Never declare width/height that do not match the actual file.

### 7. Cache-bust to force a re-unfurl

**Rule:** When you change an OG image or the meta tags and the old preview keeps showing, append `?v=2` (or any new query string) to `og:url` and the share URL.

**Why:** iMessage, WhatsApp, Slack, and Discord cache OG metadata **per exact URL**, often for days. They do not honor `Cache-Control` from your origin for this — only a different URL invalidates. Most have no public "force re-fetch" button.

**How to apply:** For one-off shares, just paste `https://example.com/page?v=2`. For a permanent re-unfurl on a marketing page, bump a version in the canonical URL and add a redirect from the old. For Facebook/LinkedIn, use their debuggers (see references/unfurler-matrix.md).

## The minimum viable ``

```html

Page Title — site-name

```

For `og:type=article` add `article:published_time`, `article:modified_time`, `article:author`, and one `article:tag` per tag.

## Character limits (truncate at render, not source)

| Field          | Limit | Notes                                                              |
| -------------- | ----- | ------------------------------------------------------------------ |
| ``      | 60    | Including site-name suffix. Truncate at word boundary with ellipsis. |
| `description`  | 155   | Same string for ``, `og:description`, `twitter:description`. |
| `og:image:alt` | 420   | Practically: one short sentence.                                   |

Store the long version in your SEO registry; truncate inside the `MetaTags` component.

## GEO (Generative Engine Optimization) — the bonus layer

Generative engines (Perplexity, ChatGPT search, Claude, Gemini, Google AI Overviews) consume the same `` metadata as social unfurlers, plus a few extras. The OpenGraph rules above already cover ~80% of GEO. Add these:

1. **JSON-LD `Article` schema** for blog/changelog/long-form pages. `@context: https://schema.org`, `@type: Article` (or `BlogPosting`, `NewsArticle`), `headline`, `description`, `image`, `datePublished`, `dateModified`, `author`. Generative engines use this as ground truth more than they use OG tags.
2. **Clear, factual first paragraph.** The first 200 characters of body text are what gets quoted. Lead with the claim, not a hook.
3. **`` matches `` semantically.** Engines penalize divergence between them as a relevance signal.
4. **Stable canonical URLs.** Do not let GEO-indexable content live behind query-string variants. Always set ``.
5. **`robots.txt` allows AI crawlers** (or the specific ones you want to be cited by). The default is conservative — explicitly allow `GPTBot`, `ClaudeBot`, `PerplexityBot`, `Google-Extended` if you want to be cited.

For the deeper schema.org patterns, defer to the page-type-specific spec when one exists in `context-v/` of the project. For broader on-page and technical SEO concerns that also feed GEO (canonicals, sitemaps, breadcrumbs, anti-patterns), see [references/seo-best-practices.md](references/seo-best-practices.md).

## Sitemap & robots.txt — proper search-engine discoverability

Every Lossless Astro site that wants to be indexed (every splash, every marketing site, anything other than gated decks) ships `sitemap-index.xml` + `sitemap-0.xml` plus `robots.txt`. The Astro team maintains an official integration that handles 95% of this for free.

### The integration

[`@astrojs/sitemap`](https://docs.astro.build/en/guides/integrations-guide/sitemap/) — official, maintained by Astro Core. Install + add to integrations array. It walks every static route Astro emits at build time and writes `dist/sitemap-index.xml` + `dist/sitemap-0.xml` with absolute URLs derived from the `site` + `base` config you already have.

```bash
pnpm add @astrojs/sitemap
# or, on bun-based splashes (memopop):
bun add @astrojs/sitemap
```

```js
// astro.config.mjs
import sitemap from '@astrojs/sitemap';

export default defineConfig({
  site: 'https://lossless-group.github.io',
  base: '/your-repo/',
  integrations: [
    sitemap({
      filter: (page) =>
        !page.includes('/llms.txt') &&
        !page.includes('/llms-full.txt') &&
        !page.endsWith('/404/') &&
        !page.endsWith('/404'),
    }),
  ],
});
```

### Hard rules (Why + How to apply)

#### Rule SM-1: Filter non-HTML routes out of the sitemap

**Why:** sitemaps are for HTML pages search engines can rank and serve to humans. The `/llms.txt` and `/llms-full.txt` endpoints are markdown for LLM consumers — they're not pages a Google searcher should land on. Likewise `/404` is a status page, not content. Including them pollutes the index, can confuse crawlers, and (in the case of `/llms-full.txt` at multiple MB) wastes crawl budget.

**How to apply:** every Lossless splash uses the filter shown above. Copy it verbatim. If your site adds new non-HTML endpoints, extend the filter — don't remove existing exclusions.

#### Rule SM-2: Ship `public/robots.txt` with an absolute `Sitemap:` line

**Why:** robots.txt is the canonical discovery path. Crawlers fetch `/robots.txt` first and follow the `Sitemap:` directive to find the sitemap. Without it, you depend on the crawler stumbling onto `/sitemap-index.xml` by convention — which works for Google but isn't guaranteed elsewhere.

**How to apply:** create `public/robots.txt`:

```
User-agent: *
Allow: /

Sitemap: https://your-host/your-base/sitemap-index.xml
```

The `Sitemap:` URL must be **absolute** (with protocol and full path). If your splash is path-deployed under GitHub Pages, the path must be in the URL — e.g., `https://lossless-group.github.io/content-farm/sitemap-index.xml`. Astro copies `public/` files through to `dist/` unchanged, so robots.txt deploys to `/robots.txt` at the host root.

For path-deployed splashes the file *also* deploys at `/your-base/robots.txt` — which is fine. Crawlers always check the host root first.

#### Rule SM-3: Add `` to BaseLayout's ``

**Why:** small, harmless, helps tools that prefer in-document discovery (some site auditors, some custom crawlers). Costs nothing.

**How to apply:**

```astro

```

Place it near the favicon `` tag — both are root-relative resource hints that belong together.

#### Rule SM-4: One integration call, one filter, no `customPages`

**Why:** the integration auto-discovers every page Astro emits. We don't have hand-rolled routes that need explicit registration. Adding `customPages` is a smell — it usually means someone is trying to fix a routing problem in the wrong place.

**How to apply:** never pass `customPages`. If a page isn't in the sitemap, it's because (a) the filter is excluding it, (b) it's not actually being rendered, or (c) it's an API/endpoint route and shouldn't be there. Diagnose at the source.

### Counting URLs (the gotcha)

Astro emits **single-line minified XML**. `grep -c '' dist/sitemap-0.xml` returns `1` because it counts *lines*, not occurrences. Use:

```bash
grep -o '' dist/sitemap-0.xml | wc -l
```

This catches everyone the first time. Add it to whatever verification recipe lives near the splash.

### Verification

```bash
# 1. Sitemap exists and is non-empty
ls -lh dist/sitemap-*.xml

# 2. URL count matches expectations (one per HTML page, minus filtered routes)
grep -o '' dist/sitemap-0.xml | wc -l

# 3. Filtered routes are absent
grep -E 'llms\.txt|llms-full\.txt|/404/' dist/sitemap-0.xml
# (corpus content with "llms" or "404" in slugs is fine — the filter operates on
# the deployed page URL, not on slug substrings within other URLs)

# 4. robots.txt has the absolute Sitemap: line
cat dist/robots.txt

# 5. Head tag is present
grep 'rel="sitemap"' dist/index.html
```

### Reference implementations

All five Lossless splashes ship this pattern as of 2026-05-09:

| Splash | URLs | Base |
|---|---|---|
| `ai-labs/context-vigilance-kit/splash` | 463 | `/context-vigilance-kit/` |
| `astro-knots/splash` | 162 | `/astro-knots/` |
| `content-farm/splash` | 95 | `/content-farm/` |
| `ai-labs/memopop-ai/apps/memopop-site` | 128 | `/memopop-ai/` |
| `lfm/splash` | 17 | `/lossless-flavored-markdown-package/` |

For the full porting recipe with copy-paste-ready config, robots.txt, and head tags, see [references/sitemap-implementation.md](references/sitemap-implementation.md).

## llms.txt — give the model the corpus directly

The [llms.txt standard](https://llmstxt.org/) is the GEO concentric ring beyond meta tags and JSON-LD: a small markdown file at the host root that lists (or contains) the machine-readable corpus of a site, ready for LLM ingest in one fetch. Where OpenGraph optimizes for the social unfurler, llms.txt optimizes for the model.

### When to ship it

- Any site with a substantive content collection — corpus, changelog, docs, blog — that you'd want models to cite or learn from
- Splash pages that aggregate child-repo content (memopop, content-farm, lfm, context-vigilance-kit)
- Documentation sites
- Marketing splashes with substantial long-form content (case studies, blog posts)

Skip it for pure UI-shell sites with no content (a single landing page), one-shot fundraise decks, and slide-only sites.

### Two files (the second is optional but encouraged)

1. **`/llms.txt`** — markdown link index. Sections per content collection, each entry as `- [Title](url): description`. Small (10s–100s of KB).
2. **`/llms-full.txt`** — concatenated raw markdown bodies of every entry, with metadata headers (title, source, canonical URL, last-modified). The preferred ingest target — avoids per-page HTTP overhead and HTML noise. Large (multi-MB on content-heavy sites).

### Hard rules (Why + How to apply)

#### Rule LLM-1: Source of truth for prose lives in markdown, not TypeScript

**Why:** Voice and framing are human-edited, not dev-edited. Putting "Treat context with the same vigilance as code…" inside a `.ts` template literal forces a developer-flavored review process for what should be a copy edit. Splash maintainers will avoid the file. The text rots.

**

…

## Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

- **Author:** [lossless-group](https://github.com/lossless-group)
- **Source:** [lossless-group/lossless-agent-skills](https://github.com/lossless-group/lossless-agent-skills)
- **License:** MIT

Install and usage instructions live in the source repository linked above.

## Pricing

- **Free** — Free

## Security capabilities

Automated source analysis of v0.1.0 — what this tool can access:

- **Network access:** yes
- **Filesystem access:** no
- **Shell / process execution:** no
- **Environment & secrets:** yes
- **Dynamic code execution:** no

*"Yes" means the capability is present in the source — more access means more to trust, not that it is unsafe.*


## Versions

- **0.1.0** — security scan: passed — Imported from the upstream source.

## Links

- Listing page: https://agentstack.voostack.com/l/skill-lossless-group-lossless-agent-skills-open-graph-share-seo-geo
- Seller: https://agentstack.voostack.com/s/lossless-group
- Browse the marketplace: https://agentstack.voostack.com/browse

---
Listed on AgentStack — the marketplace for AI agent skills and MCP servers. Every listing is security-reviewed. Creators keep 70%.
