AgentStack
SKILL verified MIT Self-run

Personal Genomics Analysis

skill-zocomputer-skills-personal-genomics-analysis · by zocomputer

Analyzes your 23andMe genetic data to identify health risks, drug sensitivities, ancestry, and traits, then generates a personalized report

No reviews yet
0 installs
15 views
0.0% view→install

Install

$ agentstack add skill-zocomputer-skills-personal-genomics-analysis

✓ scanned · ✓ verified — works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access Used
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

Are you the author of Personal Genomics Analysis? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Personal Genomics Analysis Pipeline

How to Use This Prompt

This prompt transforms raw 23andMe genome data into a comprehensive, queryable genomics knowledge base with clinical annotations and personalized health insights.

Prerequisites:

  • 23andMe raw genome data file (V3, V4, or V5 format, typically \~600k-900k SNPs)
  • The file should be in tab-delimited format with columns: rsid, chromosome, position, genotype
  • Approximately 10-15 minutes of processing time
  • \~5GB of disk space for databases

What This Pipeline Creates:

  1. genome.db - Your personal genome database (631k SNPs, \~50MB)
  2. clinvar.db - Clinical variant database (3.8M variants, 472k pathogenic, \~600MB)
  3. pharmgkb.db - Pharmacogenomics database (drug-gene interactions)
  4. Analysis scripts - Reusable Python tools for querying your genome
  5. Visualizations - Ancestry charts, trait summaries
  6. Personalized report - Comprehensive markdown document with all findings

What You'll Discover:

  • 🧬 Ancestry composition - Estimated genetic ancestry from key markers
  • 💊 Pharmacogenomics - How you metabolize common drugs (critical for medical records!)
  • 🧠 Cognitive/behavioral traits - COMT, OXTR, BDNF variants
  • 💪 Athletic performance - ACTN3, ACE genes
  • ❤️ Health risks - Diabetes, cardiovascular, Alzheimer's markers
  • 🍽️ Dietary genetics - Lactose tolerance, caffeine metabolism, taste perception
  • 👁️ Physical traits - Eye color, earwax type, hair characteristics
  • 🚨 Pathogenic variants - Screened against 472k known disease variants

Warning: This analysis is for informational and research purposes only. It is not a substitute for professional genetic counseling or medical advice. Critical findings (especially pharmacogenomics) should be shared with your healthcare provider.

Database Schemas

1. genome.db (Your Personal Genome)

CREATE TABLE snps (
    rsid TEXT PRIMARY KEY,
    chromosome TEXT NOT NULL,
    position INTEGER NOT NULL,
    genotype TEXT NOT NULL
);

CREATE INDEX idx_chromosome ON snps(chromosome);
CREATE INDEX idx_position ON snps(position);
CREATE INDEX idx_chrom_pos ON snps(chromosome, position);

Fields:

  • rsid: SNP identifier (e.g., "rs12913832")
  • chromosome: Chromosome number (1-22, X, Y, MT)
  • position: Base pair position on chromosome
  • genotype: Your two alleles (e.g., "AA", "AG", "GG")

Example queries:

-- Find a specific SNP
SELECT * FROM snps WHERE rsid = 'rs12913832';

-- Count SNPs per chromosome
SELECT chromosome, COUNT(*) as count
FROM snps
GROUP BY chromosome
ORDER BY CAST(chromosome AS INTEGER);

-- Find all SNPs in a gene region (e.g., APOE)
SELECT * FROM snps
WHERE chromosome = '19'
AND CAST(position AS INTEGER) BETWEEN 45409039 AND 45412650;

2. clinvar.db (Clinical Variants)

CREATE TABLE clinvar_variants (
    rsid TEXT PRIMARY KEY,
    chromosome TEXT,
    position INTEGER,
    ref_allele TEXT,
    alt_allele TEXT,
    clinical_significance TEXT,
    disease_name TEXT,
    review_status TEXT
);

CREATE INDEX idx_clinvar_significance ON clinvar_variants(clinical_significance);
CREATE INDEX idx_clinvar_chrom_pos ON clinvar_variants(chromosome, position);

Fields:

  • clinical_significance: Pathogenic, Likely_pathogenic, Benign, etc.
  • disease_name: Associated condition
  • review_status: Evidence quality (e.g., "practice guideline")

3. pharmgkb.db (Pharmacogenomics)

CREATE TABLE pharmgkb_variants (
    rsid TEXT PRIMARY KEY,
    gene TEXT,
    drug TEXT,
    impact TEXT,
    recommendation TEXT,
    evidence_level TEXT,
    pmid TEXT
);

Evidence levels:

  • 1A: High level of evidence (clinical guidelines)
  • 1B: Moderate level of evidence
  • 2A/2B: Lower levels of evidence
  • 3/4: Preliminary evidence

Pipeline Procedure

Step 1: Verify Input File

Locate the 23andMe raw data file. Typical format:

# rsid chromosome position genotype

rs12564807 1 734462 AA
rs3131972 1 752721 GG

Validation checks:

  1. File exists and is readable
  2. Tab-delimited format
  3. Contains rsid, chromosome, position, genotype columns
  4. At least 500,000 SNPs present
  5. Genotypes are valid (AA, AG, GG, AT, etc.)

Step 2: Create Personal Genome Database

Create a Python script to parse the 23andMe file and populate genome.db:

#!/usr/bin/env python3
import sqlite3
import logging
from pathlib import Path

logging.basicConfig(level=logging.INFO, format="%(asctime)sZ %(levelname)s %(message)s")

HEALTH_DIR = Path("")
GENOME_FILE = HEALTH_DIR / ".txt"
DB_FILE = HEALTH_DIR / "genome.db"

def create_database():
    conn = sqlite3.connect(DB_FILE)
    cursor = conn.cursor()

    cursor.execute('''
        CREATE TABLE IF NOT EXISTS snps (
            rsid TEXT PRIMARY KEY,
            chromosome TEXT NOT NULL,
            position INTEGER NOT NULL,
            genotype TEXT NOT NULL
        )
    ''')

    cursor.execute('CREATE INDEX IF NOT EXISTS idx_chromosome ON snps(chromosome)')
    cursor.execute('CREATE INDEX IF NOT EXISTS idx_position ON snps(position)')
    cursor.execute('CREATE INDEX IF NOT EXISTS idx_chrom_pos ON snps(chromosome, position)')

    conn.commit()
    return conn

def parse_and_insert(conn):
    cursor = conn.cursor()
    batch = []
    batch_size = 10000
    total_count = 0

    with open(GENOME_FILE, 'r') as f:
        for line in f:
            line = line.strip()
            if not line or line.startswith('#'):
                continue

            parts = line.split('\t')
            if len(parts) != 4:
                continue

            rsid, chrom, pos, geno = parts

            if not rsid.startswith('rs') and not rsid.startswith('i'):
                continue

            batch.append((rsid, chrom, pos, geno))

            if len(batch) >= batch_size:
                cursor.executemany(
                    "INSERT OR REPLACE INTO snps (rsid, chromosome, position, genotype) VALUES (?, ?, ?, ?)",
                    batch
                )
                total_count += len(batch)
                logging.info(f"Inserted {total_count} SNPs...")
                batch = []

        if batch:
            cursor.executemany(
                "INSERT OR REPLACE INTO snps (rsid, chromosome, position, genotype) VALUES (?, ?, ?, ?)",
                batch
            )
            total_count += len(batch)

    conn.commit()
    logging.info(f"Total SNPs in database: {total_count}")

def main():
    conn = create_database()
    parse_and_insert(conn)
    conn.close()
    logging.info(f"Database created: {DB_FILE.absolute()}")

if __name__ == "__main__":
    main()

Expected output: genome.db with 500k-900k SNPs (depending on 23andMe chip version)

Step 3: Download and Index ClinVar Database

ClinVar contains all known clinically-relevant genetic variants. Download and parse:

import asyncio
import aiohttp
import gzip
import sqlite3

CLINVAR_URL = "https://ftp.ncbi.nlm.nih.gov/pub/clinvar/tab_delimited/variant_summary.txt.gz"
CLINVAR_DB = Path("") / "clinvar.db"

async def setup_clinvar():
    # Download ClinVar (~50MB compressed, ~5GB uncompressed)
    logging.info("Downloading ClinVar database (~50MB)...")

    async with aiohttp.ClientSession() as session:
        async with session.get(CLINVAR_URL) as response:
            with open(CLINVAR_DB.parent / "variant_summary.txt.gz", 'wb') as f:
                async for chunk in response.content.iter_chunked(8192):
                    f.write(chunk)

    # Decompress and parse
    logging.info("Decompressing and parsing ClinVar...")
    conn = sqlite3.connect(CLINVAR_DB)
    cursor = conn.cursor()

    cursor.execute('''
        CREATE TABLE IF NOT EXISTS clinvar_variants (
            rsid TEXT PRIMARY KEY,
            chromosome TEXT,
            position INTEGER,
            ref_allele TEXT,
            alt_allele TEXT,
            clinical_significance TEXT,
            disease_name TEXT,
            review_status TEXT
        )
    ''')

    batch = []
    with gzip.open(CLINVAR_DB.parent / "variant_summary.txt.gz", 'rt') as f:
        header = f.readline()

        for line in f:
            parts = line.strip().split('\t')
            if len(parts) = 10000:
                cursor.executemany('''
                    INSERT OR REPLACE INTO clinvar_variants
                    VALUES (?, ?, ?, ?, ?, ?, ?, ?)
                ''', batch)
                conn.commit()
                batch = []

        if batch:
            cursor.executemany('''
                INSERT OR REPLACE INTO clinvar_variants
                VALUES (?, ?, ?, ?, ?, ?, ?, ?)
            ''', batch)
            conn.commit()

    conn.close()
    logging.info("ClinVar database created")

Note: This step takes 5-10 minutes and creates a \~600MB database.

Step 4: Create PharmGKB Database

Pharmacogenomics variants that affect drug metabolism:

def setup_pharmgkb():
    conn = sqlite3.connect(PHARMGKB_DB)
    cursor = conn.cursor()

    cursor.execute('''
        CREATE TABLE IF NOT EXISTS pharmgkb_variants (
            rsid TEXT PRIMARY KEY,
            gene TEXT,
            drug TEXT,
            impact TEXT,
            recommendation TEXT,
            evidence_level TEXT,
            pmid TEXT
        )
    ''')

    # Key pharmacogenomics markers
    variants = [
        ('rs4244285', 'CYP2C19', 'Clopidogrel', 'Poor metabolizer', 'Reduced efficacy', '1A', '23250844'),
        ('rs9923231', 'VKORC1', 'Warfarin', 'Increased sensitivity', 'Lower dose required', '1A', '22617227'),
        ('rs1799853', 'CYP2C9', 'Warfarin', 'Reduced metabolism', 'Lower dose required', '1A', '22617227'),
        ('rs776746', 'CYP3A5', 'Tacrolimus', 'Poor metabolizer', 'Higher drug levels', '1A', '22378157'),
        ('rs1801133', 'MTHFR', 'Methotrexate', 'Reduced activity', 'Toxicity risk', '2A', '21270786'),
        ('rs4680', 'COMT', 'Opioids', 'Pain sensitivity', 'Affects response', '2B', '16331281'),
        ('rs1799971', 'OPRM1', 'Opioids', 'Receptor binding', 'Affects requirement', '2A', '17363983'),
        ('rs12248560', 'CYP2C19', 'PPIs,Clopidogrel', 'Rapid metabolizer', 'Altered response', '1A', '23250844'),
    ]

    cursor.executemany('''
        INSERT OR REPLACE INTO pharmgkb_variants VALUES (?, ?, ?, ?, ?, ?, ?)
    ''', variants)

    conn.commit()
    conn.close()

Step 5: Query Your Genome Against Clinical Databases

Cross-reference your SNPs with ClinVar and PharmGKB:

def query_personal_genome():
    genome_conn = sqlite3.connect(GENOME_DB)
    clinvar_conn = sqlite3.connect(CLINVAR_DB)
    pharmgkb_conn = sqlite3.connect(PHARMGKB_DB)

    # PharmGKB findings
    logging.info("🔬 PHARMACOGENOMICS FINDINGS:")
    cursor = pharmgkb_conn.execute("SELECT * FROM pharmgkb_variants")

    for row in cursor:
        rsid, gene, drug, impact, rec, evidence, pmid = row
        result = genome_conn.execute("SELECT genotype FROM snps WHERE rsid = ?", (rsid,)).fetchone()

        if result:
            genotype = result[0]
            logging.info(f"\n📍 {rsid} ({gene})")
            logging.info(f"   Your genotype: {genotype}")
            logging.info(f"   Affects: {drug}")
            logging.info(f"   Impact: {impact}")
            logging.info(f"   Clinical: {rec}")
            logging.info(f"   Evidence: Level {evidence} | PMID: {pmid}")

    # ClinVar pathogenic variants
    logging.info("\n\n🧬 CLINVAR FINDINGS (Pathogenic/Likely Pathogenic):")

    genome_cursor = genome_conn.execute("SELECT rsid FROM snps")
    my_rsids = {row[0] for row in genome_cursor}

    clinvar_cursor = clinvar_conn.execute("""
        SELECT rsid, disease_name, clinical_significance, review_status
        FROM clinvar_variants
        WHERE clinical_significance LIKE '%Pathogenic%'
        OR clinical_significance LIKE '%Likely_pathogenic%'
    """)

    pathogenic_found = []
    for row in clinvar_cursor:
        rsid, disease, sig, review = row
        if rsid in my_rsids:
            pathogenic_found.append((rsid, disease, sig, review))

    if pathogenic_found:
        for rsid, disease, sig, review in pathogenic_found:
            genotype = genome_conn.execute("SELECT genotype FROM snps WHERE rsid = ?", (rsid,)).fetchone()[0]
            logging.info(f"\n⚠️  {rsid}")
            logging.info(f"   Genotype: {genotype}")
            logging.info(f"   Disease: {disease}")
            logging.info(f"   Significance: {sig}")
            logging.info(f"   Review: {review}")
    else:
        logging.info("\n✅ No pathogenic variants found in ClinVar database")

    genome_conn.close()
    clinvar_conn.close()
    pharmgkb_conn.close()

Step 6: Trait Analysis

Query specific SNPs for traits, ancestry, and health markers:

Key trait markers to query:

  • rs12913832 (HERC2) - Eye color
  • rs17822931 (ABCC11) - Earwax type, body odor
  • rs4988235 (LCT) - Lactose tolerance
  • rs762551 (CYP1A2) - Caffeine metabolism
  • rs1815739 (ACTN3) - Athletic performance
  • rs713598 (TAS2R38) - Bitter taste perception
  • rs4680 (COMT) - "Warrior vs Worrier" gene
  • rs7903146 (TCF7L2) - Type 2 diabetes risk
  • rs1333049 (CDKN2A/B) - Coronary artery disease
  • rs429358 + rs7412 (APOE) - Alzheimer's risk

Ancestry-informative markers:

  • rs12913832, rs1426654, rs16891982 - Pigmentation
  • rs885479 (MC1R region) - African ancestry
  • rs1800407 (OCA2) - Eye/hair color, ancestry
  • rs3827760 (EDAR) - East Asian ancestry

Step 7: Generate Visualizations

Create ancestry pie charts using matplotlib:

import matplotlib.pyplot as plt

# Calculate ancestry percentages from markers
ancestry_data = {
    'African': 37.6,
    'East Asian': 24.7,
    'South Asian': 22.7,
    'European': 15.1
}

plt.figure(figsize=(10, 6))
plt.pie(ancestry_data.values(), labels=ancestry_data.keys(), autopct='%1.1f%%')
plt.title("Estimated Ancestry Composition")
plt.savefig("ancestry_visualization.png")

Step 8: Create Personalized Health Report

Generate a comprehensive markdown report:

def create_personalized_report():
    output_file = HEALTH_DIR / "personalized_genomics_report.md"

    with open(output_file, 'w') as f:
        f.write("# Personalized Genomics Report\n\n")

        f.write("## Executive Summary\n\n")
        f.write("### 🚨 Critical Findings\n\n")
        # Include pharmacogenomics, health risks

        f.write("### ✅ Protective Factors\n\n")
        # Include beneficial variants

        f.write("## Pharmacogenomics Summary\n\n")
        # Table of drug-gene interactions

        f.write("## Dietary Recommendations\n\n")
        # Based on metabolic genes

        f.write("## Exercise Recommendations\n\n")
        # Based on athletic performance genes

        f.write("## Data Sources\n\n")
        f.write("- Genome: 23andMe (XXX,XXX SNPs)\n")
        f.write("- ClinVar: 3.88M variants\n")
        f.write("- PharmGKB: Clinical pharmacogenomics\n")

    return output_file

Step 9: Create Reusable Query Scripts

Build Python scripts that let you (or anyone else) query the genome easily:

#!/usr/bin/env python3
"""Query a specific SNP from the genome database"""
import sqlite3
import sys

def query_snp(rsid):
    conn = sqlite3.connect("/path/to/genome.db")
    result = conn.execute("SELECT * FROM snps WHERE rsid = ?", (rsid,)).fetchone()

    if result:
        rsid, chrom, pos, geno = result
        print(f"SNP: {rsid}")
        print(f"Location: chr{chrom}:{pos}")
        print(f"Genotype: {geno}")
    else:
        print(f"SNP {rsid} not found in database")

    conn.close()

if __name__ == "__main__":
    if len(sys.argv) != 2:
        print("Usage: python query_snp.py ")
        sys.exit(1)

    query_snp(sys.argv[1])

Expected Output

After running this pipeline, you should have:

  1. Databases:
  • genome.db (50-100MB)
  • clinvar.db (\~600

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet — be the first.

Versions

  • v0.1.0 Imported from the upstream source.