Install
$ agentstack add skill-thirdwatch-dev-scraping-skills-ecommerce-product-scraping ✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ● Network access Used
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
Verified badge
Passed review? Show it. Paste this badge into your README, it links to the public security report.
Reliability & compatibility
Declared compatibility
Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.
We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.
How agent discovery & health will work →About
E-commerce Product Scraping
Routes product-data tasks to the right scraper or build path. The goal is structured product rows — title, brand, price, MRP/list price, discount, rating, variants, image, URL — at the lowest cost per result.
How storefronts expose product data
Most e-commerce sites embed full product JSON in the page, so you rarely need to parse HTML:
- SSR / embedded JSON —
__NEXT_DATA__(Next.js),__NUXT__,window.__INITIAL_STATE__, or site-specific blobs (Myntra's__myx). Full structured catalog with no DOM parsing. - Internal search/PDP APIs — open DevTools → Network → XHR while browsing; listings and product pages almost always fetch JSON you can call directly. Find the endpoint the page calls for itself rather than parsing HTML.
- Stealth-browser-only — a few storefronts sit behind Akamai or DataDome (Nykaa, Meesho, Tata Cliq) and serve a skeleton to HTTP clients. These need a hardened browser; intercept the XHR or walk product anchors rather than fighting hashed CSS-module class names.
Key fields to capture: title, brand, price, mrp/list price, discount, rating, variants/sizes, availability, image, url. MRP + discount together are what make a row useful for price/MAP monitoring.
Ready-made scrapers
Skip the anti-bot fight — these are maintained and billed pay-per-result.
| Target | Scraper | From | Notes | |--------|---------|------|-------| | Amazon | Amazon Products | $0.002/result | 18+ country domains, ASIN/price/rating | | Flipkart | Flipkart Products | $0.003/result | India, price/rating/seller | | AliExpress | AliExpress Products | $0.003/result | price/sales count/ratings | | Myntra | Myntra | $0.002/result | India fashion, sizes + variants | | Nykaa | Nykaa | $0.005/result | India beauty | | Meesho | Meesho | $0.005/result | India social commerce | | Snapdeal | Snapdeal | $0.002/result | India marketplace | | Noon.com | Noon | $0.002/result | UAE/KSA/Egypt | | AJIO | AJIO | $0.002/result | India fashion | | FirstCry | FirstCry | $0.002/result | India baby/kids | | Tata Cliq | Tata Cliq | $0.005/result | India marketplace | | Shopify Store | Shopify Store Products | $0.001/result | any Shopify store, full variants/SKU/price | | Shopify Reviews | Shopify Reviews | $0.002/result | review count + rating across widgets |
Run one
Each scraper takes a JSON input and returns product rows. Run synchronously from the CLI:
curl -X POST "https://api.apify.com/v2/acts/thirdwatch~amazon-product-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"queries": ["wireless earbuds"],
"maxResults": 25
}'
Get a free token at console.apify.com. Exact input fields are on each actor's Store page — input keys differ per site (search term vs. category URL vs. product URL).
Build your own
No scraper for your target, or you need to own it? Start with the engineering skills in this collection:
web-scraping-playbook— the build-vs-buy decision and the cost-first technique ladder (HTTP → TLS spoof → stealth browser).anti-bot-scraping— concrete bypasses for Akamai, DataDome, Cloudflare, PerimeterX (the protections in front of Nykaa, Meesho, Tata Cliq).apify-actor-builder— package it as a deployable, monetizable Apify Actor.
Maintained by Thirdwatch. 70+ ready-made scrapers on the Apify Store.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: thirdwatch-dev
- Source: thirdwatch-dev/scraping-skills
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet, be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.