AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified Unlicense Self-run

Accelerating Python

skill-yale-som-hpc-claude-code-marketplace-accelerating-python · by yale-som-hpc

Make slow or memory-bound Python fast on the Yale SOM HPC cluster — profile, push work into DuckDB/Polars/Arrow, vectorize/Numba, then parallelize or use a GPU. Also covers data larger than a job's RAM (Parquet, query engines, sampling). TRIGGER when a Python job on the cluster is slow, CPU-bound, or memory-heavy, when a dataset on /gpfs does not fit a job's RAM, or when weighing parallelism or G…

No reviews yet
0 installs
9 views
0.0% view→install

Install

$ agentstack add skill-yale-som-hpc-claude-code-marketplace-accelerating-python

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-yale-som-hpc-claude-code-marketplace-accelerating-python)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Accelerating Python? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Accelerating Python

Rule: profile first; query before loading; then pick the smallest acceleration that matches the bottleneck. Most "slow Python" and "data too big" problems are solved before you reach for multiprocessing or a GPU.

Order of operations

  1. Measure the slow part — don't guess.
  2. Reduce data with filters/projections before loading (push them into the file reader).
  3. Use better engines: DuckDB, Polars, Arrow, NumPy — they stream and push down.
  4. Vectorize simple array/dataframe operations.
  5. Numba for tight numeric loops that don't vectorize cleanly.
  6. Multiprocessing for CPU-bound Python — see [parallel python](../parallel-python/SKILL.md).
  7. GPU only for GPU-shaped work — see [using GPUs](../using-gpus/SKILL.md).

Do not start at step 6 if the real bottleneck is CSV parsing, GPFS metadata, a database query, or network latency.

Quick profiler

import cProfile
import pstats

with cProfile.Profile() as profiler:
    main()

stats = pstats.Stats(profiler).sort_stats("cumtime")
stats.print_stats(25)

For a running job, py-spy dump --pid PID is often more useful (if installed).

Store data in a columnar format

  • Parquet for tabular data, Arrow datasets for partitioned data, HDF5 for array-like data.
  • Avoid thousands of CSVs, repeated CSV parsing, and Excel as an intermediate format.

One reused Parquet beats 10k CSVs for both speed and GPFS metadata health.

Query before loading

The single biggest win for both speed and memory: filter and project in the reader so the full table never enters Python. DuckDB (SQL) and Polars (lazy) both push predicates/projections into the Parquet scan and stream the result — pick whichever idiom matches your code.

DuckDB — SQL over Parquet with pushdown:

import duckdb

duckdb.sql("""
COPY (
  SELECT gvkey, fyear, sales
  FROM read_parquet('/gpfs/project/myproject/data/raw/panel/*.parquet')
  WHERE fyear >= 2010
) TO '/gpfs/project/myproject/data/derived/sales.parquet'
(FORMAT PARQUET)
""")

Polars — lazy expressions with a streaming sink:

import polars as pl

(
    pl.scan_parquet("/gpfs/project/myproject/data/raw/panel/*.parquet")
    .filter(pl.col("fyear") >= 2010)
    .select("gvkey", "fyear", "sales")
    .sink_parquet("/gpfs/project/myproject/data/derived/sales.parquet")
)

Neither materializes the full table in Python. Either is usually faster and simpler than a hand-written Python loop.

Installing the CLI tools

Neither duckdb nor qsv is module-loadable on the cluster (module spider duckdb / qsv both return "Unable to find"). Install them as static binaries in ~/.local/bin — see [installing software](../installing-software/SKILL.md#user-binaries) for PATH setup:

mkdir -p ~/.local/bin
# DuckDB CLI — single static binary, latest release
curl -L https://github.com/duckdb/duckdb/releases/latest/download/duckdb_cli-linux-amd64.zip -o /tmp/duckdb.zip
unzip -o /tmp/duckdb.zip -d ~/.local/bin && rm /tmp/duckdb.zip
duckdb --version
# qsv — pick the musl prebuild (x86_64-unknown-linux-musl) from
# https://github.com/jqnatividad/qsv/releases ; same drop-in install.

Inside a Python project you also get DuckDB via uv add duckdb — that's the library (import duckdb); the command-line duckdb is a separate binary, both backed by the same engine. CLI for ad-hoc inspection, library inside scripts.

Quick command-line inspection

Look before you script:

duckdb -c "DESCRIBE SELECT * FROM 'data.csv'"
duckdb -c "SELECT category, COUNT(*) FROM 'data.csv' GROUP BY category ORDER BY 2 DESC"
duckdb -c "COPY (SELECT * FROM 'data.csv') TO 'data.parquet' (FORMAT PARQUET)"   # convert reusable extracts
qsv stats data.csv          # fast CSV triage

Sample first

Before a full run, debug code and estimate memory on a sample:

sample = duckdb.sql("""
SELECT * FROM read_parquet('/gpfs/project/myproject/data/raw/panel/*.parquet')
USING SAMPLE 10000 ROWS
""").df()

Numba

A JIT compiler for a subset of Python/NumPy. Reach for it when tight loops over NumPy arrays dominate and interpreter overhead is the bottleneck — not as a general "make Python fast" knob.

import numpy as np
from numba import njit

@njit(cache=True)            # cache=True reuses compiled code across runs
def scale(arr):
    out = np.empty_like(arr)
    for i in range(arr.shape[0]):
        out[i] = arr[i] * 2
    return out

Parallel loops need parallel=True; parallelize the outer loop:

from numba import njit, prange

@njit(parallel=True, cache=True)
def add_one(arr):
    for i in prange(arr.shape[0]):
        arr[i] += 1

Race warning: parallel prange writing to a shared array (e.g. counts[idx[i]] += 1) races. Use per-worker accumulators and reduce, or np.add.at(counts, idx, 1) outside Numba (thread-safe). And Numba compiles on first call — benchmark the second call to exclude compile time.

SQLite for small caches

Local caches and lookup tables; enable WAL when multiple readers use the cache:

import sqlite3
conn = sqlite3.connect("/gpfs/project/myproject/cache/lookup.db")
conn.execute("PRAGMA journal_mode=WAL")

One connection per thread/process; SQLite serializes writes — keep one writer path.

One file per array task

For job arrays, write output/task_0001.parquet, …, then combine — don't mutate one big Parquet. The full append/atomic-write patterns live in [using the filesystem](../using-the-filesystem/SKILL.md).

pl.scan_parquet("output/task_*.parquet").sink_parquet("output/all_results.parquet")

Other languages

Writing R? arrow + dplyr over open_dataset() gives the same pushdown — see [coding in R](../coding-in-r/SKILL.md) and [running R](../running-r/SKILL.md).

Joins need real keys

Large-data tools do not make invalid joins valid. Use documented identifiers or official link tables (e.g. CRSP-Compustat CCM linking); finance-specific linking belongs in a finance-data skill.

Checklist

  • [ ] The bottleneck is measured, not guessed.
  • [ ] Filters/projections are pushed into the scan/query before materializing.
  • [ ] Raw data is stored once in Parquet; files inspected with DuckDB/qsv before big scripts.
  • [ ] Code is tested on a sample before full data; memory checked after the test job.
  • [ ] DuckDB/Polars/NumPy are tried before Python loops; Numba uses @njit(cache=True) and avoids shared-write races.
  • [ ] Multiprocessing/GPU are used only when single-process optimization isn't enough.
  • [ ] Outputs are Parquet/HDF5, not thousands of tiny CSVs; large merges use documented keys.

Further reading

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.