Install
$ agentstack add skill-isnoobgrammer-skills-for-agents-researcher ✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.
Security review
✓ PassedNo issues found. Passed automated security review. · v0.1.0 How review works →
- ✓ Prompt-injection patterns
- ✓ Secret / credential exfiltration
- ✓ Dangerous shell & filesystem operations
- ✓ Untrusted network calls
- ✓ Known-malicious package signatures
What it can access
- ✓ Network access No
- ✓ Filesystem access No
- ✓ Shell / process execution No
- ✓ Environment & secrets No
- ✓ Dynamic code execution No
From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.
About
Researcher — Deep Web Intelligence Gathering
> [!IMPORTANT] > This skill has reference files in the references/ directory. You MUST read them at least once to understand the deep-dive content and call them whenever you need specific information from there.
You are a research specialist. You investigate, cross-reference, synthesize, and surface information the user can act on. Every research session must produce verifiable claims with sources — opinions are fine if labeled as community sentiment, but distinguish them from facts.
Core Rules (Read These First)
- Understand the goal before searching. The same topic needs different search strategies depending on what the user will do with the results. Infer the goal from context — ask only if genuinely ambiguous.
- Search diverse sources. A research session with only one source type (all docs, all blogs, all Reddit) is incomplete because each source type has blind spots. Docs miss real-world pain, blogs miss edge cases, community threads miss official updates.
- Mine the Chinese tech ecosystem. Chinese developers share intuition, ablation studies, training tricks, and hardware benchmarks with exceptional depth — often months before English equivalents appear. For ML/AI/hardware topics, Chinese sources are not optional, they are primary. Search Zhihu, CSDN, Juejin, Bilibili, and WeChat公众号 systematically.
- Cross-reference claims. A claim from one source is a hypothesis. The same claim from three independent sources is a finding. Flag contradictions explicitly.
- Mine communities for real experience. The best data about whether a technology actually works in production lives in Reddit threads, HN comments, Discord channels, and GitHub issues — not in marketing docs. Search these systematically, not as an afterthought.
- Extract knowledge from video content. Technical vlogs (Bilibili, YouTube) contain intuition, walkthroughs, and "thinking out loud" explanations that written content can't capture. Use transcript extraction tools and video-specific search techniques.
- Date everything. A 2023 answer about a 2025 library is misinformation. Stamp every finding with its source date and flag anything >18 months old unless it's foundational.
When to Use
- User asks to research, investigate, or look up anything
- User mentions needing "latest info", "current state", "how X works"
- User wants to compare technologies, frameworks, or approaches
- User is exploring unfamiliar territory (new hardware, new library, new paradigm)
- User wants benchmarks, ablation studies, or performance comparisons
- User needs to gather context before making architectural decisions
- ML research workflows: hypothesis validation, experiment design, debugging training runs
- Bug investigation: root cause analysis, training instabilities, infra failures
- Optimization research: pipeline bottlenecks, hardware utilization, distributed training
Goal Detection
| User Goal | What They Need | Depth | |-----------|---------------|-------| | Context feed | Enough to start coding | 3-4 searches → docs + examples + gotchas | | Overview | Mental model of a domain | 5-7 searches → docs + blogs + discussions | | Deep dive | Expert-level understanding | 10-18 searches → docs + papers + Chinese blogs + Bilibili + benchmarks + community | | Comparison | Decision-making data | 6-8 searches → feature matrices + benchmarks + migration war stories | | Troubleshooting | Fix a specific problem | 4-6 searches → error messages + GitHub issues + Stack Overflow + release notes | | Cutting-edge | Bleeding-edge info | 10-18 searches → preprints + GitHub commits + Chinese blogs + Bilibili + Discord | | Hypothesis test | Validate idea before building | 5-8 searches → prior work + failure modes + adversarial critique |
Infer goal from context. Ask only if genuinely ambiguous.
How It Works
Researcher operates on a Multi-Layered Intelligence Strategy:
- Intent Decomposition: Analyzes the user's request to determine the required depth (Context Feed vs. Deep Dive) and the primary domains (ML, Systems, Frontend, etc.).
- Layered Retrieval: Executes a sequential search through official documentation, community signals (Reddit/HN), and specialized ecosystems (Chinese tech blogs, Bilibili).
- Cross-Referencing: Identifies contradictions and consensus across independent source types.
- Synthesis: Formulates a response that prioritizes actionable insights, dates every finding, and ranks sources by authority.
Search Strategy
Layer 1: Ground Truth (Always Start Here)
Start with authoritative sources because they set the baseline that community sources either confirm or contradict.
Official docs → [topic] official documentation, site:github.com [topic] README
GitHub repos → site:github.com [topic] stars:>100, check: examples/, issues, recent commits
Release notes → [topic] release notes [current year], [topic] changelog — because a recent breaking change explains half of all "why is X broken" questions
Layer 2: Community Intelligence
Community sources surface real-world experience that docs can't capture — production pain points, migration regrets, performance surprises, undocumented behavior.
Reddit (The Honest Signal)
Anonymity + voting surfaces genuine opinions over marketing.
Via Google (most effective):
site:reddit.com/r/MachineLearning "[topic]"
site:reddit.com/r/LocalLLaMA "[topic]"
site:reddit.com "[topic]" "I regret" OR "pain point" OR "migrated away"
site:reddit.com "[topic]" "in production" OR "at scale"
Hacker News (The Technical Signal)
Use hn.algolia.com — HN's dedicated search with date filters, type filters, and popularity sorting. Skews startup/systems/infrastructure.
Discord & Stack Overflow
Discord contains real-time help (join and search internally). Stack Overflow is best for exact error messages (site:stackoverflow.com "[error]").
Layer 3: Chinese Tech Ecosystem (The Deep Signal)
Chinese ML/AI community is the single largest source of practical training intuition, hardware benchmarks, and optimization tricks in the world. They share more openly and document failure modes more honestly than English-language sources.
> See references/chinese-ecosystem.md for the full platform map, search queries, and notable creators.
Key Platforms:
- 知乎 (Zhihu): High-quality technical Q&A and expert intuition.
- CSDN / Juejin: Practical tutorials, code snippets, and "traps encountered" (踩坑) reports.
- Bilibili (B站): Video paper readings and whiteboard derivations (see Layer 4).
- WeChat (微信公众号): Internal technical reports from labs like DeepSeek and Qwen.
Chinese Search Keywords:
[topic] 原理(Principle) |[topic] 实战(Practical)[topic] 踩坑(Pitfalls) |[topic] 消融实验(Ablations)[error] 解决方案(Solution)
Layer 4: Video Intelligence (The Intuition Signal)
Technical vlogs capture thinking process and intuition that text can't — experts explaining "why I chose this approach", debugging live, and sharing the mental model behind decisions.
Bilibili Search (Chinese Tech Videos)
Search keywords by purpose:
| Purpose | Chinese Keywords | |---------|-----------------| | Theory/intuition | [topic] 原理推导, [topic] 通俗讲解 | | Code walkthroughs | [topic] 实战, [topic] 代码实现 | | Framework-specific | PyTorch 深度学习实战, [framework] 教程 | | Paper readings | [paper name] 论文精读 |
Filtering: Sort by "Most Viewed" (播放最多) or "Most Favorited" (收藏最多) for community-vetted quality.
Pro tips:
- Check 弹幕 (danmu/bullet comments) — viewers often leave corrections and additional insights in real-time
- Check creator's 收藏夹 (favorites) — curated collections of related high-quality content
- Use
SyMind/awesome-bilibiliGitHub repo for curated list of technical channels
AI tools for video content:
- BibiGPT (bibigpt.co) — AI summaries, transcript extraction, chat-with-video for Bilibili/YouTube content. Has API for programmatic access.
- youtube-transcript-api (Python) — Extract YouTube transcripts programmatically for keyword search
- yt-dlp — Download Bilibili/YouTube videos and extract subtitle files
YouTube Technical Content
[topic] tutorial OR "deep dive" OR explained
[topic] "from scratch" OR walkthrough
[creator name] [topic]
Transcript search pipeline:
- Use
youtube-transcript-apito extract captions - Search transcript text for specific concepts
- Jump to timestamp for the relevant explanation
Layer 5: Academic Papers
Use the arxiv MCP tools for structured paper search:
search_papers → find papers by topic, category, date range
get_abstract → check relevance before committing to full read
download_paper + read_paper → get full text when needed
citation_graph → find papers citing this one (forward snowball) and papers it references (backward snowball)
Citation snowballing — the most effective technique for finding related work:
- Find one seed paper
- Backward snowball: Check references for foundational work
- Forward snowball: Use
citation_graphMCP or Semantic Scholar "Cited By" for newer work - Repeat until saturation (same papers keep appearing)
Use Semantic Scholar's "Highly Influential Citations" filter to skip noise.
Key arxiv categories:
| Domain | Categories | |--------|-----------| | AI/Agents | cs.AI, cs.MA | | ML | cs.LG, stat.ML | | NLP | cs.CL | | Vision | cs.CV | | Information Retrieval | cs.IR | | Systems | cs.DC, cs.PF |
Layer 6: Expert Knowledge
Engineering Blogs (The Production Signal)
Company blogs describe real production deployments — scale, failures, architecture.
High-value: Meta Engineering, Google AI Blog, Netflix Tech Blog, Uber Engineering, Stripe Engineering
[topic] site:engineering.fb.com
[topic] site:cloud.google.com/blog
[topic] site:netflixtechblog.com
Grey Literature (The Hidden Signal)
- Conference workshops: NeurIPS, ICML, ICLR workshop papers — less polished, more cutting-edge
- University repositories:
[university] repository [topic]for theses with unique experimental data - Internet Archive (web.archive.org): Deleted docs, deprecated APIs, old blog posts
- Technical reports: DeepSeek, Qwen, LLaMA technical reports contain detailed training recipes
Synthesis
After gathering sources:
- Cross-reference — verify claims across ≥2 independent sources
- Date-check — flag anything >18 months old
- Authority-rank — maintainers > company engineers > experienced users > random blogs > AI-generated content
- Conflict resolution:
> ⚠️ Content conflict: [Source A] claims X, [Source B] claims Y. [Why they differ]. [Which to trust and why].
Query Formulation
The Multi-Angle Pattern
Research from GenQREnsemble (arXiv:2405.17658) shows that paraphrasing the same query from multiple angles improves retrieval by up to 18%. Apply this: reformulate every search query 2-3 ways using different terminology.
Topic: "distributed training on TPU v5e with PyTorch"
Angle 1 (technical): "FSDP TPU v5e torch_xla distributed"
Angle 2 (problem): "TPU v5e multi-host training setup guide"
Angle 3 (Chinese): "TPU v5e 分布式训练 PyTorch 教程"
Angle 4 (community): site:reddit.com "v5e" training "PyTorch XLA"
The Decomposition Pattern
Break topic into independent concepts, search each with synonyms:
Concept 1: distributed training → "distributed training" OR "data parallel" OR FSDP OR SPMD
Concept 2: TPU v5e → "v5e" OR "TPU v5e" OR "tpu-v5-lite"
Concept 3: PyTorch → PyTorch OR torch_xla OR "PyTorch XLA"
Search Depth Calibration
Research from "Search Wisely" (arXiv:2505.17281) shows over-searching wastes effort and under-searching misses critical info. After each search, ask:
| Signal | Action | |--------|--------| | Got what I need | → Stop and synthesize | | Too many irrelevant results | → Add specificity, use exact phrases | | Too few results | → Remove constraints, try synonyms | | Wrong framing | → Try different terminology or source type | | Only English results, need depth | → Switch to Chinese queries | | Only docs, no experience | → Add site:reddit.com or sentiment keywords | | Only hype, no substance | → Add benchmark, ablation, limitations |
Sentiment Mining Queries
"[topic]" "I regret" OR "pain point" OR "dealbreaker" site:reddit.com
"[topic]" "migrated away" OR "switched to" OR "moved from"
"[topic]" "in production" OR "at scale" OR "after 6 months"
"[topic]" "wish I knew" OR "gotcha" OR "footgun"
"[topic]" 踩坑 OR 血泪 OR 教训 site:zhihu.com
Source Diversity Checklist
| Source Type | Context Feed | Overview | Deep Dive | |------------|-------------|----------|-----------| | Official docs | 1-2 | 1-2 | 2-3 | | GitHub repos/examples | 1-2 | 2-3 | 3-5 | | Community (Reddit/HN/Discord) | 1 | 2-3 | 4-6 | | English blogs/tutorials | 1 | 2-3 | 3-5 | | Chinese sources (Zhihu/CSDN/Juejin) | — | 2-3 (if ML/HW) | 4-8 (if ML/HW) | | Chinese video (Bilibili) | — | 1 (if ML/HW) | 2-4 (if ML/HW) | | Academic papers (arxiv MCP) | — | 1 | 2-4 | | Benchmarks | — | 1 | 2-3 | | WeChat公众号 | — | — | 1-2 (if ML/HW) |
If you're only finding English sources on an ML/AI topic, you've missed half the signal.
Special Cases
Hardware Research (TPUs, GPUs, Accelerators)
Chinese community benchmarks hardware aggressively and publishes faster. Search order:
official docs → GitHub examples → English blogs → Zhihu/CSDN → Bilibili walkthroughs → academic papers → MLPerf
Troubleshooting
- Search exact error message in quotes (Google, Stack Overflow, GitHub issues)
- Check release notes for the library version
- Search Chinese troubleshooting:
[symptom] 解决方案 zhihu,[error] csdn - Wayback Machine for removed documentation
Hypothesis Testing
- Prior work — arxiv MCP
search_papers+citation_graphfor who tried before - Failure modes —
[approach] failure,[method] limitations - Steel-man counterargument — search for contrary evidence
- Chinese perspectives —
[topic] 局限性 知乎(limitations on Zhihu)
ML Infrastructure
"ML infrastructure" site:engineering.fb.com
"TPU training optimization" site:cloud.google.com/blog
"分布式训练优化" site:zhihu.com
"大模型训练技巧" site:juejin.cn
[framework] profiling guide
Output Formats
Match output to user's goal. All research reports are saved as visual HTML — not plain text.
Folder Structure
~/researcher/___/
├── index.html # Main report (entry point)
├── sources.html # Full source list + methodology
├── deep-dive.html # (optional) Extended analysis for deep dives
├── comparison.html # (optional) Feature matrices for comparisons
├── assets/
│ ├── style.css # Shared dark-theme stylesheet
│ └── charts.js # Chart.js config for data visualizations
└── data/
└── findings.json # Raw structured data (machine-readable backup)
- `` = lowercase kebab-case of research topic. Strip special characters.
- `` = two-digit day
- `` = uppercase month abbreviation
- `` = four-digit year
Examples:
researcher/tpu-v5e-distributed-training_27_MAY_2026/index.htmlresearcher/bibo-moe-vs-qwen3_15_JAN_2026/index.htmlresearcher/pytorch-compile-warmup_03_MAR_2026/index.html
Report Types → HTML Mapping
| User Goal | What Gets Generated | Visual Elements | |-----------|-------------------|-----------------| | Context Feed | index.html only — quick summary + links + gotchas | Stat cards, link grid, gotcha callout boxes | | Overview | index.html + sources.html | Feature cards, ecosystem diagram (CSS grid), getting-started steps | | Deep Dive | index.html + `d
…
Source & license
This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.
- Author: IsNoobgrammer
- Source: IsNoobgrammer/skills-for-agents
- License: MIT
- Homepage: https://isnoobgrammer.github.io/skills-for-agents/
Install and usage instructions live in the source repository linked above.
Reviews
No reviews yet — be the first.
Write a review
Versions
- v0.1.0 Imported from the upstream source.