AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Benchmark

skill-code-saurabh-openskills-benchmark · by CODE-SAURABH

Performance benchmarking skill — baseline and compare page load times, Core Web Vitals, bundle sizes, and API response times. Use when the user wants to measure performance before and after a change, catch performance regressions in a PR, establish a performance baseline, or understand what is making their app slow.

No reviews yet
0 installs
17 views
0.0% view→install

Install

$ agentstack add skill-code-saurabh-openskills-benchmark

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-code-saurabh-openskills-benchmark)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Benchmark? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Benchmark

Benchmarking without a baseline is noise. A benchmark that can't be reproduced is a guess. Every benchmark in this skill produces a number, a method, and a comparison — not an impression.

Performance is a feature. Regressions ship silently. This skill makes regressions visible before they reach users.


Benchmark Principles

  • Measure before you optimize. An optimization without a before measurement is an assumption. Always baseline first.
  • Reproduce exactly. A benchmark is only valid if another engineer can reproduce it with the same setup. Document the environment, the command, and the conditions.
  • Medians lie. Use percentiles. P50 hides the users who are suffering. Always report P75, P95, and P99 alongside the median.
  • Synthetic vs real-user data. Lighthouse in a lab tells you what's possible. RUM (Real User Monitoring) tells you what's happening. Use both.
  • One change at a time. If two things change between benchmarks, you can't attribute the delta to either.
  • A regression is a regression. A 10% slowdown on a fast page is still a regression. Set thresholds and enforce them.

Step 0: Establish the Baseline

Before any optimization or PR review, establish the baseline. Without it, "after" numbers are meaningless.

# Run Lighthouse baseline (3 runs, take median)
npx lighthouse https://your-app.com/key-page \
  --output json \
  --output-path ./baseline.json \
  --chrome-flags="--headless" \
  --throttling-method=simulate \
  --preset=desktop

# Run 3 times and average — single runs are noisy
for i in 1 2 3; do
  npx lighthouse https://your-app.com/key-page \
    --output json \
    --output-path "./baseline-run-$i.json" \
    --chrome-flags="--headless"
done

Record in benchmark-baseline.md:

## Baseline — [Page/Endpoint] — [Date] — [Commit SHA]

| Metric | Run 1 | Run 2 | Run 3 | Median |
|--------|-------|-------|-------|--------|
| LCP    |       |       |       |        |
| INP    |       |       |       |        |
| CLS    |       |       |       |        |
| FCP    |       |       |       |        |
| TTFB   |       |       |       |        |
| Total JS bundle | | | | |
| Total CSS bundle | | | | |
| Network requests | | | | |

**Environment:** macOS M2 / 6-core CPU throttle / Fast 3G simulation
**Tool:** Lighthouse 12.x / Chrome 124
**URL:** https://your-app.com/key-page

Core Web Vitals — Reference

| Metric | Full Name | Measures | Good | Needs Improvement | Poor | |--------|-----------|----------|------|-------------------|------| | LCP | Largest Contentful Paint | Loading | ≤2.5s | 2.5–4.0s | >4.0s | | INP | Interaction to Next Paint | Interactivity | ≤200ms | 200–500ms | >500ms | | CLS | Cumulative Layout Shift | Visual stability | ≤0.1 | 0.1–0.25 | >0.25 | | FCP | First Contentful Paint | Loading start | ≤1.8s | 1.8–3.0s | >3.0s | | TTFB | Time to First Byte | Server response | ≤800ms | 800ms–1.8s | >1.8s |


Frontend Benchmarking

Lighthouse (automated, reproducible)

# Single page audit
npx lighthouse https://your-app.com \
  --output=json,html \
  --output-path=./lighthouse-report \
  --chrome-flags="--headless --no-sandbox"

# Key metrics to extract from JSON
cat lighthouse-report.report.json | jq '{
  lcp: .audits["largest-contentful-paint"].numericValue,
  inp: .audits["interaction-to-next-paint"].numericValue,
  cls: .audits["cumulative-layout-shift"].numericValue,
  fcp: .audits["first-contentful-paint"].numericValue,
  ttfb: .audits["server-response-time"].numericValue,
  score: .categories.performance.score
}'

Bundle Size Analysis

# Next.js
npx next build 2>&1 | grep -A 50 "Route (app)"

# Webpack Bundle Analyzer
npm install --save-dev webpack-bundle-analyzer
# Add to webpack config, then:
npx webpack --profile --json > stats.json
npx webpack-bundle-analyzer stats.json

# Source map explorer (any bundler)
npm install --save-dev source-map-explorer
npx source-map-explorer 'build/static/js/*.js'

# Check for duplicate packages
npx duplicate-package-checker-webpack-plugin

Bundle Size Regression Check

# Before (on main branch)
git checkout main && npm run build
du -sh .next/static/chunks/*.js | sort -rh | head -10 > bundle-before.txt

# After (on feature branch)
git checkout feature-branch && npm run build
du -sh .next/static/chunks/*.js | sort -rh | head -10 > bundle-after.txt

diff bundle-before.txt bundle-after.txt

API Benchmarking

Response Time (k6)

// k6-script.js
import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  stages: [
    { duration: '30s', target: 10 },   // ramp up
    { duration: '1m',  target: 10 },   // steady state
    { duration: '10s', target: 0 },    // ramp down
  ],
  thresholds: {
    http_req_duration: ['p(95) r.status === 200,
    'response time  r.timings.duration max)max=$1} END {printf "min=%.3fs avg=%.3fs max=%.3fs\n", min, sum/NR, max}'

Database Query Benchmarking

-- Explain analyze a slow query
EXPLAIN (ANALYZE, BUFFERS, FORMAT TEXT)
SELECT u.*, o.total
FROM users u
JOIN orders o ON o.user_id = u.id
WHERE u.created_at > NOW() - INTERVAL '30 days'
ORDER BY o.total DESC
LIMIT 100;

-- Check for sequential scans on large tables (should be index scans)
-- Look for: Seq Scan on large_table — this is the red flag

-- Find slow queries in PostgreSQL
SELECT query,
       calls,
       mean_exec_time::numeric(10,2) AS mean_ms,
       max_exec_time::numeric(10,2)  AS max_ms,
       total_exec_time::numeric(10,2) AS total_ms
FROM pg_stat_statements
ORDER BY mean_exec_time DESC
LIMIT 20;

Before / After Comparison Format

Always produce this table when benchmarking a change:

## Benchmark Results — [Feature/Change] — [Date]

**Commit before:** abc1234  
**Commit after:**  def5678  
**Environment:** [describe exactly]  
**Tool:** [Lighthouse 12 / k6 1.0 / custom]

### Core Web Vitals (median of 3 runs)

| Metric | Before | After | Delta | Verdict |
|--------|--------|-------|-------|---------|
| LCP    | 2.8s   | 2.1s  | -25%  | ✅ Improved |
| INP    | 180ms  | 210ms | +17%  | ⚠️ Regressed |
| CLS    | 0.05   | 0.05  |   0%  | ✅ Neutral |
| Perf score | 72 | 81  | +9pts | ✅ Improved |

### Bundle Size

| Asset | Before | After | Delta | Verdict |
|-------|--------|-------|-------|---------|
| main.js | 284KB | 301KB | +6%  | ⚠️ Regressed |
| vendor.js | 412KB | 398KB | -3% | ✅ Improved |
| CSS total | 48KB | 48KB |  0%  | ✅ Neutral |

### API Response Time (p50 / p95 / p99)

| Endpoint | Before | After | Delta | Verdict |
|----------|--------|-------|-------|---------|
| GET /api/users | 45/120/280ms | 38/95/210ms | -25% | ✅ Improved |
| POST /api/orders | 120/380/950ms | 125/390/960ms | +1% | ✅ Neutral |

### Overall Verdict
🟢 SHIP — performance improved overall. INP regression is within acceptable range (180→210ms, still 4.0s |
| INP | +15% | +30% or >500ms |
| CLS | any increase | >0.1 absolute |
| JS bundle | +5% | +15% or +50KB |
| API P95 | +10% | +25% or >1s |
| Lighthouse score | -3pts | -10pts |

---

## CI Integration (Lighthouse CI)

```yaml
# .github/workflows/lighthouse.yml
name: Lighthouse CI
on: [pull_request]

jobs:
  lighthouse:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
      - run: npm ci && npm run build
      - run: npm run start &
      - uses: treosh/lighthouse-ci-action@v11
        with:
          urls: |
            http://localhost:3000/
            http://localhost:3000/dashboard
          budgetPath: ./lighthouse-budget.json
          uploadArtifacts: true

# lighthouse-budget.json
[{
  "path": "/*",
  "timings": [
    { "metric": "largest-contentful-paint", "budget": 2500 },
    { "metric": "cumulative-layout-shift",  "budget": 0.1 }
  ],
  "resourceSizes": [
    { "resourceType": "script", "budget": 400 },
    { "resourceType": "total",  "budget": 800 }
  ]
}]

Bundled Resource

Use python scripts/compare_metrics.py baseline.json candidate.json to compare flat JSON metric files in a repeatable way. Read [metric-contract.md](./references/metric-contract.md) first to define units and regression thresholds; do not compare incompatible measurements.

Definition of Done — Benchmark

  • [ ] Baseline established before any change (commit SHA recorded)
  • [ ] Same environment used for before and after measurements
  • [ ] Minimum 3 runs taken, median reported (not single run)
  • [ ] All Core Web Vitals measured (LCP, INP, CLS, FCP, TTFB)
  • [ ] Bundle sizes measured before and after
  • [ ] API P50/P95/P99 measured for changed endpoints
  • [ ] Before/after comparison table produced
  • [ ] Every metric has a verdict: Improved / Regressed / Neutral
  • [ ] Any regression explained (is it acceptable? why?)
  • [ ] Overall ship/hold verdict stated with rationale
  • [ ] Results committed to benchmark-results/ for historical comparison

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.