AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Seo Technical

skill-thibaultbm-claude-seo-geo-seo-technical · by Thibaultbm

Find and fix what stops a site from being indexed, fast, and visible to AI crawlers. Input: a site or specific pages. Output: fixes for indexation, robots.txt, sitemaps, IndexNow, Core Web Vitals (LCP/INP/CLS), JS rendering, AI crawler access (GPTBot, ClaudeBot, PerplexityBot, and more), canonicals, hreflang, HTTPS, and migrations. Repo reference for AI crawler control. Use for a technical audit,…

No reviews yet
0 installs
33 views
0.0% view→install

Install

$ agentstack add skill-thibaultbm-claude-seo-geo-seo-technical

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access Used
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-thibaultbm-claude-seo-geo-seo-technical)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
2mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Seo Technical? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Technical SEO: crawl, index, render

Technical SEO decides whether content can be fetched, parsed, indexed, and trusted by Google, Bing, and AI answer engines (ChatGPT, Claude, Perplexity, Gemini). In 2026 the discipline split in two: Googlebot renders JavaScript and tolerates slow pages, AI crawlers do neither. Treat the raw HTML response as the product.

Work in this order: indexation, rendering, speed. A fast page that is not indexed earns nothing. An indexed page whose content only appears after client-side JavaScript runs is invisible to every AI engine except Google's. Every threshold below is labeled either official or measured (with a source URL) or field heuristic from 115+ agency audits.

Company knowledge first (Obsidian)

If the working environment contains an Obsidian vault or any local knowledge base (a folder of .md notes, often with a .obsidian directory), read the relevant notes before acting: brand and product facts, target keywords, competitors, and the SEO action log of what was already tried. Ground every recommendation in that context instead of asking the user for facts the vault already holds. At the end of the session, append the actions taken to the vault's SEO action log so the next session starts informed. Vault structure, read-first and write-back protocols: the obsidian-brain skill.

When to use this skill

Use it for:

  • Full technical audits and pre-sales triage of any site
  • "Why is my site (or page) not indexed" investigations
  • robots.txt, meta robots, sitemap, canonical, and hreflang work
  • Core Web Vitals (LCP, INP, CLS), PageSpeed, image weight
  • JavaScript rendering decisions: SPA versus SSR versus SSG
  • AI crawler access: this skill owns references/ai-crawlers.md, the canonical crawler reference for the whole skill set
  • llms.txt questions (see the honest verdict section)
  • Site migrations, redesigns, domain changes
  • Hacked-site cleanup (injected spam links)

Hand off neighboring problems:

  • Orphan pages and link architecture: use the seo-internal-linking skill
  • JSON-LD and structured data: use the seo-schema-markup skill
  • Writing content so AI engines quote it: use the geo-visibility skill
  • Measuring AI citations and AI referral traffic: use the geo-tracking skill
  • One-pass scored audit of a whole site: use the seo-geo-audit skill

Workflow

Step 0: Secure measurement access

Before any recommendation, confirm both:

  1. Google Search Console is verified (domain property preferred) and you can actually open it.
  2. GA4, or an equivalent analytics tool, is installed and collecting.

Field finding from 115+ agency audits: a large share of audited sites had neither installed, or had them installed with nobody able to access the accounts. Without GSC you cannot see coverage, queries, or manual actions, and every later step degrades into guessing. Set both up first, then add Bing Webmaster Tools: it imports GSC properties in a few clicks and matters in 2026 because Bing's index feeds ChatGPT search (https://yoast.com/chatgpt-search/).

For a fast automated first pass, run the bundled audit script from the seo-geo-audit skill (scripts/seo_audit.py), which checks robots.txt AI bot rules, sitemap, and on-page basics. Use its output to direct the manual work below.

Step 1: Read the indexation state

In GSC, open Indexing > Pages and compute the ratio of indexed to submitted pages.

  • 75 percent or more indexed (for example 3,000 indexed of 4,000 submitted) is healthy on a mature site (field heuristic).
  • Below 50 percent, stop and investigate before anything else: duplication, thin pages, crawl traps, leftover noindex, or rendering failures.

Then act on the gaps:

  • Priority pages not indexed: inspect each with the URL Inspection tool, fix the reported cause, then click Request Indexing. The quota is roughly 10 to 12 requests per day per property, so spend it on money pages, not the archive.
  • Bulk indexation, not page by page: ignore the "index 10 URLs per day" ritual some platforms (Wix and others) push. Manual one-by-one submission scales to nothing. Submit the full sitemap in GSC once and let Google discover every URL from it; reserve manual Request Indexing for a handful of priority pages (field heuristic from 115+ agency audits).
  • Scaled and programmatic pages, the indexation condition: thousands of generated pages (one per city, per combination) only get indexed when each page carries genuinely unique content. The same block copied across all of them is the classic non-indexation cause: Google treats near-duplicate templates as one page and drops the rest. Uniqueness per page is the price of admission for scaled content, not an optimization (field heuristic from 115+ agency audits).
  • 404s with history: any URL that previously earned traffic or backlinks and now returns 404 gets a 301 to the closest equivalent page. Never blanket-redirect everything to the homepage: Google treats irrelevant mass redirects as soft 404s.
  • Orphan pages: target zero. Detection and fixes belong to the seo-internal-linking skill.

Measurement artifact (September 2025): Google removed the num=100 results parameter. Many properties saw desktop impressions collapse and average position improve overnight because rank-tracking bots stopped loading 100-result pages. Check clicks before declaring a loss: clicks typically stayed flat. Never present this artifact to a client as a traffic drop (https://searchengineland.com/google-num100-impact-data-462231).

Step 2: Crawl controls: robots.txt, sitemaps, meta robots

robots.txt:

  • Must contain a Sitemap: line pointing at the sitemap index.
  • Must not contain Disallow: / (shipped staging configs cause this more often than expected; check it first on any sudden deindexation).
  • AI bot rules: decide with the GEO layer below and references/ai-crawlers.md.

Sitemap.xml:

  • Segment by content type (sitemap-pages.xml, sitemap-posts.xml, sitemap-products.xml) under one index. Why: GSC reports coverage per sitemap, so segmentation reveals which template fails to get indexed.
  • Keep lastmod honest: update it only when content meaningfully changes. Google states it ignores lastmod when it is uniformly fresh or inaccurate.
  • Do not ping Google's sitemap endpoint: deprecated in June 2023, it returns 404. Submit through GSC and the robots.txt Sitemap: line (https://developers.google.com/search/blog/2023/06/sitemaps-lastmod-ping).
  • Set up IndexNow for Bing: instant push of new and updated URLs, built into many CMS plugins and Cloudflare. Bing is a primary gateway into ChatGPT search, which makes Bing Webmaster Tools standard equipment in 2026 (https://yoast.com/chatgpt-search/).

Discovery for ultra-fresh topics: when the site publishes time-sensitive news, add a dedicated News sitemap plus an RSS or Atom feed so Google can pick up items within minutes; an RSS feed is also the fastest way to source breaking topics worth covering. This channel only fits genuinely fresh, newsworthy content, not the evergreen archive.

Meta robots versus robots.txt, the distinction behind most "why is this still indexed" tickets:

| Goal | Correct tool | Why | |---|---|---| | Keep a page out of the index | meta robots noindex, page stays crawlable | Google must fetch the page to see the noindex | | Keep crawlers out of infinite or private URL spaces | robots.txt Disallow | Saves crawl budget; does NOT deindex | | Both on the same URL | Never combine them | Disallow prevents Google from ever seeing the noindex; the URL can stay indexed from external links |

A disallowed page can appear in results as a URL-only listing. If something must disappear, allow the crawl, serve noindex (or 404/410), and use the Removals tool for urgent cases.

Crawl budget (sites above roughly 10,000 URLs):

  • Faceted navigation, calendar archives, and internal search results generate near-infinite URL spaces. Disallow them in robots.txt and keep them out of sitemaps.
  • Read GSC Settings > Crawl stats: spikes on parameter URLs mean budget burned on noise instead of money pages.
  • Below a few thousand URLs, crawl budget is never the bottleneck; do not invoice work on it (field heuristic).

Step 3: JavaScript rendering, the number one technical check of 2026

No AI crawler executes JavaScript. The Vercel and MERJ study of 500M+ crawler fetches found that GPTBot, ClaudeBot, PerplexityBot, and Meta-ExternalAgent sometimes download JavaScript files (11.5 percent of ChatGPT fetches, 23.8 percent of Claude fetches) but execute none of it. Googlebot is the only major crawler that renders (https://vercel.com/blog/the-rise-of-the-ai-crawler and https://www.gsqi.com/marketing-blog/ai-search-javascript-rendering/).

Consequence map:

| Engine | Sees client-side rendered content | |---|---| | Google search, AI Overviews, AI Mode, Gemini grounding | Yes (Googlebot renders) | | ChatGPT (search and training) | No | | Claude (search and training) | No | | Perplexity | No | | Bing and Copilot | Rendering is deferred and inconsistent; treat as no |

The test costs one minute. Fetch the raw HTML and search for a sentence that should be on the page:

curl -sL https://example.com/page/ | grep -c "exact sentence from the page"

Zero matches means AI engines see an empty shell. Cross-check with view-source in a browser (not DevTools Elements, which shows the rendered DOM) and with GSC URL Inspection for what Googlebot renders.

Decision rule: any page that should rank or be cited must serve its content in the initial HTML response. SSR or SSG is mandatory; client-side rendering is acceptable only for logged-in app surfaces.

Stack verdicts:

| Stack | Content in raw HTML | Verdict | |---|---|---| | WordPress, Shopify, Webflow | Yes | Hold up at volume (field experience, 115+ audits) | | Next.js, Nuxt, Astro, SvelteKit with SSR/SSG enabled | Yes | Fine, but verify template by template: hybrid apps regress silently | | React, Vue, Angular client-side SPA | No | Invisible to every AI engine; prerender or rebuild | | Framer, Lovable, and similar recent builders | Mixed | Field experience: unstable for SEO at scale (rendering, redirects, sitemap control); audit before committing a content operation to them |

Inside a stable CMS, the theme still decides the technical floor. Pick a lean, SEO-oriented e-commerce theme (clean markup, no JavaScript bloat, content in raw HTML, fast Core Web Vitals out of the box) over a heavy multi-purpose one: on Shopify, for example, the gap between a performance-focused theme and a bloated one shows up directly in LCP. Choose the domain on the same logic: for a French local business, a .fr signals geography to Google and tends to earn a better local click-through than a .com (field heuristic from 115+ agency audits). Settle the domain before launch, since changing it later means a full migration (Step 7).

Step 4: Speed and Core Web Vitals

Judge mobile first: mobile carries 80 to 90 percent of traffic on typical lead-gen and e-commerce sites (field range). Measure with the free PageSpeed Insights web interface at https://pagespeed.web.dev: no API key, no account, and it returns both field data (CrUX, what real users experienced over 28 days) and lab data. Field data wins arguments; lab data locates causes.

| Metric | Good (at p75) | Notes | |---|---|---| | LCP | Under 2.5 s | Largest Contentful Paint | | INP | Under 200 ms | Replaced FID in March 2024; roughly 43 percent of sites failed INP at the switch, making it the most commonly failed vital (https://web.dev/articles/inp) | | CLS | Under 0.1 | Layout stability | | PageSpeed score | 75+ desktop acceptable, aim higher | Field heuristic; mobile scores run lower, judge the trend |

Keep perspective: Core Web Vitals act as a tiebreaker, not a dominant ranking factor. Do not spend weeks chasing a score of 100 on a site with an indexation or rendering problem; sequence speed after Steps 1 to 3.

High-yield fixes, ordered by how often audits surface them:

  1. Images: 200 KB or less each (field heuristic), WebP or AVIF, explicit width and height attributes to reserve layout space (prevents CLS), responsive srcset.
  2. Never lazy-load above-the-fold images, least of all the LCP image; mark the LCP image fetchpriority="high".
  3. Defer render-blocking third-party scripts; load chat widgets and trackers after interaction.
  4. Text-to-HTML ratio under 10 percent signals a thin or bloated page (field heuristic): too little content or too much markup. Related: keep the DOM under roughly 1,500 nodes; Lighthouse flags excessive DOM size starting around 800 (https://developer.chrome.com/docs/lighthouse/performance/dom-size).

Step 5: Duplicates and international: canonicals and hreflang

Canonicals (the e-commerce deduplication tool):

  • Faceted navigation, sort parameters, and session IDs multiply URLs. Point every variant at the clean URL with rel=canonical.
  • A canonical is a hint, not a directive: Google ignores canonicals contradicted by stronger signals (internal links and sitemaps pointing at variants). Align all signals on the clean URL.
  • Every page self-canonicalizes. It costs nothing and absorbs parameter pollution.
  • Quick win, the homepage canonical: many sites set canonicals on every inner page but leave the homepage without one, even though it is the most linked URL and the one that collects the most parameter and tracking variants. Add a self-referencing canonical on the home page; it is a two-minute fix with outsized cleanup value (field heuristic from 115+ agency audits).

hreflang (international versions):

  • Every language or region page lists all alternates including itself, plus one x-default pointing at the language selector or default market page.
  • Reciprocity is mandatory: if the FR page references the EN page, the EN page must reference the FR page back, or Google discards the pair.
  • Prefer hreflang in sitemaps on large sites: one file to audit instead of tags scattered across templates.
  • hreflang routes users; it does not consolidate authority. Near-duplicate language variants (en-US versus en-GB) still need distinct value or consolidation.

Step 6: Security and domain integrity

  • HTTPS everywhere: every HTTP URL 301s to HTTPS in one hop, no mixed content, valid certificate.
  • Hack detection: search the raw HTML (hacks often cloak, so not just the rendered page) for injected casino, pharma, and replica anchors. Check GSC Security Issues and run a site: search for pages you do not recognize. On WordPress: clean the injection, update core and plugins, install Wordfence, rotate all credentials.
  • Unintentional geo-blocking: a firewall rule or a country block (often shipped by a WAF, a CDN security preset, or a hosting default) can make the site unreachable from entire regions, so a slice of human visitors, search bots, and AI crawlers, and your own audit, see nothing. Symptoms: timeouts or 403s from one country but not another, missing impressions from a market that should convert. Test by fetching the site through a different-country exit (a VPN, or curl from a server in the target region) and compare the status code against your local result. Whitelist Googlebot, Bingbot, and the AI crawler ranges (references/ai-crawlers.md), and lift any country block that overlaps a real audience.
  • Domain age is an asset: an old domain carries link history and trust. Never let a legacy domain expire; renew it and 301 it. Field rule: domains are cheap, history is not replaceable.

Step 7: Migrations (when applicable)

Migrations destroy more rankings than algorithm updates do (field observation across 115+ audits). Operating rule: it is easy to break rankings and very hard to build them.

  1. Keep the same domain if humanly possible, and change one thing at a time: domain, URL structure, design, or content. Stacking all four into one launch makes any later diagnosis impossible.
  2. Freeze a full URL inventory before launch: GSC performance export (16 months), GSC coverage, the sitemap, an

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.