AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified Apache-2.0 Self-run

Ai Operations

skill-harness-harness-skills-ai-operations · by harness

>-

No reviews yet
0 installs
29 views
0.0% view→install

Install

$ agentstack add skill-harness-harness-skills-ai-operations

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-harness-harness-skills-ai-operations)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
3mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Ai Operations? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

AI Operations

Configure AI-powered predictive failure analysis and intelligent alert correlation using Harness AIDA.

Instructions

Step 1: Establish Scope

Confirm the user's org, project, service, and observability stack.

Call MCP tool: harness_list
Parameters:
  resource_type: "project"
  org_id: ""

Step 2: Identify the AI Operations Task

Determine which workflow the user needs:

  1. Predictive Failure Analysis -- ML-based detection of impending failures before SLO breach
  2. Alert Correlation and Noise Reduction -- Group related alerts and suppress duplicates

Step 3: Configure Predictive Failure Analysis

Gather from the user:

  • Service name and data sources (Datadog, Prometheus, CloudWatch)
  • Prediction horizon (30 minutes, 1 hour, 4 hours, 24 hours ahead)
  • Training data period (30 days, 90 days, 6 months)
  • Model type preference (anomaly detection, time series forecasting, ensemble)

Configure failure prediction scenarios:

  1. Memory leak detection -- Flag services where memory grows above threshold per window
  2. Disk exhaustion -- Predict time-to-full and alert N hours in advance
  3. Connection pool saturation -- Alert when pool usage exceeds threshold for sustained duration
  4. Latency degradation -- Detect progressive slowdown before SLO breach
  5. Deployment-induced regression -- Correlate metric changes with recent deployments

Configure alerting:

  • Set prediction confidence threshold (suppress below threshold to reduce noise)
  • Route alerts to PagerDuty, Slack, or other channels
  • Enable auto-generated runbook suggestions using AIDA
  • Set up false positive feedback loop for model improvement

Configure data sources:

  • Metrics source (Prometheus, Datadog, CloudWatch)
  • Log source (Elasticsearch, Splunk, CloudWatch Logs)
  • Trace source (Jaeger, Datadog APM, AWS X-Ray)
  • Model retraining frequency (daily, weekly, monthly, on data drift)

Step 4: Configure Alert Correlation and Noise Reduction

Gather from the user:

  • Current alert volume (alerts/day) and target reduction percentage
  • Alerting tools in use (PagerDuty, OpsGenie, Grafana, Datadog)
  • Correlation preferences

Configure alert correlation:

  • Correlation window: group alerts fired within N minutes
  • Correlation method: topology-based, time-based, ML-based, or hybrid
  • Service dependency mapping for topology-based correlation

Configure noise reduction:

  • Deduplication: merge identical alerts across sources
  • Suppression: suppress known-noisy alerts during maintenance windows
  • Aggregation: combine N similar alerts into a single incident
  • Priority scoring: ML-based severity assignment using historical resolution data

Configure intelligent routing:

  • Route to the team that owns the affected service
  • Escalation: auto-escalate if not acknowledged within SLA
  • Context enrichment: attach recent deployments, related logs, and runbook links to alerts

Examples

  • "Set up predictive failure analysis for our payment service" -- Configure ML models to detect memory leaks, disk exhaustion, and latency degradation
  • "Reduce our alert noise by 50%" -- Configure alert correlation and deduplication to reduce daily alert volume
  • "Alert us 4 hours before disk runs out" -- Configure disk exhaustion prediction with advance warning
  • "Correlate alerts across our microservices" -- Set up topology-based alert correlation using service dependency map
  • "Auto-generate runbook suggestions for alerts" -- Enable AIDA-powered runbook recommendations

Performance Notes

  • ML models need 2-4 weeks of baseline data before predictions become reliable -- expect higher false positive rates initially.
  • Topology-based correlation requires an accurate service dependency map -- stale maps cause missed correlations.
  • Alert correlation windows should balance grouping (longer = fewer alerts) with response time (shorter = faster notification).
  • Retraining frequency should match how fast the system changes -- fast-moving services need weekly retraining.

Troubleshooting

High False Positive Rate

  • Increase the confidence threshold to suppress low-confidence predictions
  • Provide false positive feedback to improve the model
  • Check that training data period includes representative traffic patterns (weekday, weekend, peak)

Predictions Not Triggering

  • Verify data sources are connected and sending metrics
  • Check that the prediction horizon is appropriate for the failure mode
  • Ensure the model has completed initial training (2-4 weeks minimum)

Alert Correlation Missing Related Alerts

  • Increase the correlation window to capture cascading failures
  • Update the service dependency map if topology-based correlation is in use
  • Check that all alert sources are integrated (missing sources cause orphaned alerts)

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

  • Author: harness
  • Source: harness/harness-skills
  • License: Apache-2.0
  • Homepage: https://developer.harness.io/docs/platform/harness-ai/harness-skills/

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.