Install
$ agentstack add skill-arimanyus-hermes-merchant-job-scraper-pipeline ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ● Filesystem access Used
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Job Scraper + Gatekeeper Pipeline
Automated job scraping, deduplication, and gatekeeper scoring for ML/AI job hunting.
When to Use
- Cron job that scrapes AI company career pages on Greenhouse, AshbyHQ, and Lever
- Runs on a schedule (e.g. every 15m, 6h, or daily)
- Appends new jobs to the queue file
- Scores pending jobs against profile keywords
Configuration
All PII, paths, and scoring keywords live in profile.yaml at the repo root. Copy profile.yaml.example to profile.yaml and edit. Keys this skill reads:
paths.queue— queue JSON location (default~/hermes-merchant/state/queue/jobs.json)paths.scraper_log— log filescoring.required_keywords,scoring.adjacent_roles,scoring.reject_keywords,scoring.ai_companies,scoring.tiers
Minimal loader (Python):
import yaml, os
from pathlib import Path
cfg = yaml.safe_load(Path("profile.yaml").read_text())
queue_path = Path(os.path.expanduser(cfg["paths"]["queue"]))
queue_path.parent.mkdir(parents=True, exist_ok=True)
Scraping Strategy
Sources — What Works / What Doesn't
| Source | Method | Status | Notes | |--------|--------|--------|-------| | Naukri.com | browsernavigate | Blocked | Search results require JS rendering; homepage loads but listings aren't accessible. | | Indeed.com | browsernavigate | Blocked by Cloudflare | "Additional Verification Required"; cannot bypass without residential proxies. | | Together AI | Browser → https://job-boards.greenhouse.io/togetherai | Works | Greenhouse format — use the JS snippet below to extract all jobs. | | Cohere | Browser → https://jobs.ashbyhq.com/cohere | Works (~100+ roles) | AshbyHQ — largest pipeline source. Departments include Agentic Platform, Modeling, Applied-ML, Inference, Model Serving. | | Baseten | Browser → https://jobs.ashbyhq.com/baseten | Works | AshbyHQ. Many roles previously bounced on email — prefer web form. | | Modal Labs | Browser → https://www.modal.com/careers | Works (~24 roles) | AshbyHQ embedded on page — scroll to see listings. | | Cerebras | Browser → https://www.cerebras.ai/join-us → click ML dept | Works | Greenhouse-based, limited ML roles. | | Mistral AI | Browser → https://jobs.lever.co/mistral | Works | Lever ATS — do NOT add query params; go direct and filter in-page. | | LangChain | Browser → https://jobs.ashbyhq.com/langchain | Works | AshbyHQ; iframe on their own careers page — go direct. | | Pinecone | Browser → https://www.pinecone.io/careers | Works | AshbyHQ-hosted, JS console extraction works. | | Groq | Browser → https://www.groq.com/careers/ | Moved | All engineering roles now on an external ATS (Gem.com). | | Hugging Face | Browser → https://huggingface.co/jobs | Login-gated | Requires auth to view. | | Anthropic | Browser → https://job-boards.greenhouse.io/anthropic | Works partially | Greenhouse — some listings require a second fetch. | | Qdrant | Browser | Blocked | careers.qdrant.tech resolves to a private network. |
ATS Platform Taxonomy
Three ATS platforms dominate AI company career pages. Each wants a different scraping approach:
| ATS | URL Pattern | Best Method | |-----|-------------|-------------| | Greenhouse | job-boards.greenhouse.io/{company} | browser_navigate → browser_console JS snippet | | AshbyHQ | jobs.ashbyhq.com/{company} | browser_navigate → browser_snapshot | | Lever | jobs.lever.co/{company} | browser_navigate → browser_snapshot (no query params) |
Known Greenhouse: Anthropic, Together AI, Cerebras. Known AshbyHQ: Cohere, Baseten, LangChain, Modal, Pinecone. Known Lever: Mistral AI.
Greenhouse Job Extraction
Many AI companies host on Greenhouse at https://job-boards.greenhouse.io/{company}.
JavaScript snippet to extract ALL job links from a Greenhouse page (run via browser_console):
"use strict"; (() => {
const links = document.querySelectorAll('a[href*="/jobs/"]');
const seen = new Set();
return Array.from(links)
.filter(l => l.href && !seen.has(l.href) && seen.add(l.href))
.map(l => l.href + '\t' + l.textContent.trim())
.join('\n');
})()
Returns a newline-separated list of URL\tJob Title pairs. Greenhouse pages load all jobs on one page — no pagination.
web_extract is unreliable for job scraping
web_extractfrequently fails with409 BILLING_ERRORon job-heavy targets (Naukri, Indeed, Anthropic, Cohere, Mistral, Together AI, Groq, LangChain, Baseten).web_searchis the working fallback — it returns titles, descriptions, and listing URLs from search-engine indexes. Use it to discover individual job-listing URLs.- Parallel
web_searchcalls hit 409 — always run sequentially. - Fallback chain when
web_extractfails:
web_searchfirst (sequentially) to discover URLs.- If
web_extractfails (409/504), use search-result snippets as job data — they contain title, company, and location. - For career pages, use
browser_navigatedirectly to the ATS URL.
Indeed / Naukri
- Indeed: blocked by Cloudflare.
site:indeed.comweb_searchqueries sometimes return snippet data with enough info (title, company, location) to score. - Naukri: some LLM/GenAI index pages (e.g.
naukri.com/llm-engineer-jobs-16) extract successfully viaweb_extract; search pages don't.
Gatekeeper Scoring
Load thresholds and keyword lists from profile.yaml:
tiers = cfg["scoring"]["tiers"]
required = cfg["scoring"]["required_keywords"]
adjacent = cfg["scoring"]["adjacent_roles"]
reject = cfg["scoring"]["reject_keywords"]
ai_cos = set(cfg["scoring"]["ai_companies"])
Tier Thresholds
- Tier 1:
score >= tiers.tier_1(default 75) → tailor resume, apply immediately. - Tier 2: `tiers.tier2 = t["tier1"] else ("tier-2" if score >= t["tier_2"] else "tier-3")
return score, tier, matched
### AI companies use generic SWE titles
Cohere, Baseten, Modal, and friends use generic titles like "Member of Technical Staff", "Forward Deployed Engineer", or "Agentic Platform". These score low without the AI-company context bonus because the title doesn't contain explicit ML keywords.
Roles like "Forward Deployed Engineer, Agentic Platform" (Cohere) score 0 by title alone. Include `agentic`, `post-training`, `inference`, `forward deployed`, `prompt specialist`, `applied researcher`, `research engineer`, `ml researcher`, `ai researcher`, `reinforcement learning`, `rlhf`, and `agent` in either `adjacent_roles` or `required_keywords` to catch them.
### Reject keyword false positives
"Research Engineer" contains "engineer" but is NOT a reject. Reject logic should only trigger on exact matches like "java developer" or "react developer" — not any string containing "developer".
## Job Data Shape
```json
{
"url": "https://jobs.ashbyhq.com/cohere/...",
"source": "cohere",
"title": "Forward Deployed Engineer, Agentic Platform",
"company": "Cohere",
"location": "Toronto; New York",
"experience": "Not specified",
"score": 35,
"tier": "tier-2",
"matched_keywords": [],
"scraped_at": "2026-04-19T17:35:45+00:00",
"status": "pending"
}
Status Values
pending— in queue, not yet applied.applied— successfully applied via web form.applied_email— applied via email.failed— application attempt failed; includefailure_reasonandmanual_apply_url.blocked— known blocker (visa, experience mismatch, etc.).
Deduplication
Before appending, compute existing_urls = {j["url"] for j in queue} and skip URLs already present.
Known Gotchas
- f-string + dict access:
f"score={j['score']}"with nested dict access confuses older Python parsers. Use string concatenation or.format(). - Tuple unpacking: if
score_jobreturns(score, tier, matched), unpack all three at every call site. - Parallel web calls:
web_search/web_extractin parallel reliably hits409 BILLING_ERROR. Run sequentially. - LangChain Ashby: the jobs board is iframed on the careers page. Go direct to
https://jobs.ashbyhq.com/langchain— don't try to scrape the embed. - Mistral Lever: times out in browser without residential proxies on some networks. Mark blocked and move on.
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: arimanyus
- Source: arimanyus/hermes-merchant
- License: MIT
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.