AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Ai Foundry Posture Check

skill-ricmmartins-azure-sre-agent-skills-08-ai-foundry-posture · by ricmmartins

>

No reviews yet
0 installs
18 views
0.0% view→install

Install

$ agentstack add skill-ricmmartins-azure-sre-agent-skills-08-ai-foundry-posture

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access No
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-ricmmartins-azure-sre-agent-skills-08-ai-foundry-posture)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
17d ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Ai Foundry Posture Check? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

AI Foundry & OpenAI Posture Check

Purpose

Assess the security, reliability, and cost efficiency of Azure OpenAI and Microsoft Foundry deployments. Detects the most common anti-patterns that startups make when building AI-powered products — from exposed endpoints to runaway token costs.

Based on the Azure Well-Architected Framework for AI workloads and Microsoft Foundry operational best practices.

When to use this skill

  • User asks "is our OpenAI deployment secure?"
  • User asks about AI cost optimization or token consumption
  • Review before going to production with an AI feature
  • User asks about content filtering, model versions, or rate limiting
  • Periodic AI workload health check

Pre-check

Confirm with the user:

  • Which Azure OpenAI / Cognitive Services accounts to assess (or "all in subscription")
  • Whether they use PTU (Provisioned Throughput) or Standard deployments
  • Whether they have production AI workloads already live

Assessment procedure

Step 0: Discover AI resources

az cognitiveservices account list \
  --subscription  \
  --query "[?kind=='OpenAI' || kind=='AIServices'].{name:name, kind:kind, rg:resourceGroup, location:location, sku:sku.name}" \
  -o table

If no results, try:

az cognitiveservices account list \
  --subscription  \
  --query "[].{name:name, kind:kind, rg:resourceGroup, location:location}" \
  -o table

If no Cognitive Services accounts exist, report "No Azure OpenAI or AI Foundry resources found" and end assessment.

For each account found, run the following checks:


🔐 CATEGORY 1 — Security (Critical)

Check 1.1 — Managed Identity enabled (not API keys only)

az cognitiveservices account show \
  --name  --resource-group  \
  --query "{identity:identity.type, disableLocalAuth:properties.disableLocalAuth}" \
  -o json

| Finding | Severity | Score | |---------|----------|-------| | identity.type = SystemAssigned/UserAssigned AND disableLocalAuth = true | ✅ Pass | 12 pts | | identity.type set but disableLocalAuth = false | ⚠️ Partial | 6 pts | | identity.type = None or null | ❌ Fail | 0 pts |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/managed-identity

Check 1.2 — Network isolation (firewall or private endpoint)

az cognitiveservices account show \
  --name  --resource-group  \
  --query "{publicAccess:properties.publicNetworkAccess, defaultAction:properties.networkAcls.defaultAction, privateEndpoints:properties.privateEndpointConnections}" \
  -o json

| Finding | Severity | Score | |---------|----------|-------| | publicNetworkAccess = Disabled + private endpoints exist | ✅ Pass | 13 pts | | defaultAction = Deny + IP/VNet rules configured | ✅ Pass | 13 pts | | defaultAction = Allow (open to internet) | ❌ Fail | 0 pts |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/cognitive-services-virtual-networks

Check 1.3 — Content filtering (RAI policies)

az rest --method get \
  --url "https://management.azure.com/subscriptions//resourceGroups//providers/Microsoft.CognitiveServices/accounts//raiPolicies?api-version=2024-10-01"

Then check deployments have a policy assigned:

az cognitiveservices account deployment list \
  --name  --resource-group  \
  --query "[].{name:name, model:properties.model.name, raiPolicy:properties.raiPolicyName}" \
  -o table

| Finding | Severity | Score | |---------|----------|-------| | All deployments have RAI policy assigned with filters enabled | ✅ Pass | 10 pts | | Some deployments missing policy or using minimal filters | ⚠️ Partial | 5 pts | | No RAI policies or content filtering disabled | ❌ Fail | 0 pts |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/content-filters

⚙️ CATEGORY 2 — Reliability & Operations

Check 2.1 — Model version pinned and not deprecated

az cognitiveservices account deployment list \
  --name  --resource-group  \
  --query "[].{name:name, model:properties.model.name, version:properties.model.version}" \
  -o table

Cross-reference with available models:

az cognitiveservices account list-models \
  --name  --resource-group  \
  --query "[].{model:model.name, version:model.version, lifecycle:model.lifecycleStatus}" \
  -o table

| Finding | Severity | Score | |---------|----------|-------| | All deployments on GA versions with explicit version pinned | ✅ Pass | 7 pts | | Deployments on versions nearing retirement ( 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/model-retirements

Check 2.2 — Diagnostic settings enabled

# Get resource ID
RESOURCE_ID=$(az cognitiveservices account show \
  --name  --resource-group  --query "id" -o tsv)

az monitor diagnostic-settings list --resource $RESOURCE_ID -o json

| Finding | Severity | Score | |---------|----------|-------| | Diagnostic settings sending RequestResponse + Audit to Log Analytics | ✅ Pass | 7 pts | | Diagnostic settings exist but incomplete (only metrics, no logs) | ⚠️ Partial | 3 pts | | No diagnostic settings at all | ❌ Fail | 0 pts |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/monitoring

Check 2.3 — Resource Lock on production account

az lock list --resource-group  \
  --query "[?contains(id, '')].{name:name, level:level}" \
  -o table

| Finding | Severity | Score | |---------|----------|-------| | CanNotDelete lock exists on the OpenAI account | ✅ Pass | 5 pts | | Lock exists on RG but not specifically on account | ⚠️ Partial | 3 pts | | No locks | ❌ Fail | 0 pts |

> 📖 https://learn.microsoft.com/en-us/azure/azure-resource-manager/management/lock-resources

Check 2.4 — Multi-region resilience (failover capability)

Check if OpenAI accounts exist in more than one region (disaster recovery / latency optimization):

az cognitiveservices account list \
  --subscription  \
  --query "[?kind=='OpenAI' || kind=='AIServices'].{name:name, location:location, rg:resourceGroup}" \
  -o table

Group by location and count unique regions:

| Finding | Severity | Score | |---------|----------|-------| | OpenAI accounts in 2+ different regions | ✅ Pass | 5 pts | | All OpenAI accounts in a single region | ⚠️ Risk | 2 pts | | Only 1 OpenAI account total (single point of failure) | ❌ Risk | 0 pts |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/business-continuity-disaster-recovery

> Context from the field: Customers with single-region deployments face outages during regional capacity events. Multi-region + APIM load balancing (Check 4.1) is the recommended pattern for production AI workloads.

Check 2.5 — 429 throttling rate (last 24 hours)

Requires diagnostic settings sending logs to Log Analytics (Check 2.2). If diagnostics are not configured, skip and note dependency.

AzureDiagnostics
| where TimeGenerated > ago(24h)
| where ResourceProvider == "MICROSOFT.COGNITIVESERVICES"
| where Category == "RequestResponse"
| summarize 
    TotalRequests = count(),
    ThrottledRequests = countif(resultSignature_d == 429),
    ThrottleRate = round(100.0 * countif(resultSignature_d == 429) / count(), 2)
  by Resource, properties_s
| order by ThrottleRate desc

Use execute_kusto_query with the Log Analytics workspace connected to the OpenAI account's diagnostic settings.

| Finding | Severity | Score | |---------|----------|-------| | Throttle rate 10% (active capacity problem) | ❌ Fail | 0 pts | | No diagnostic data available (Check 2.2 failed) | ⚠️ Skip | 0 pts — flag dependency |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/quota

> Context from the field: Customers assume PTU eliminates throttling — it does not. Bursty traffic and concurrency limits can cause 429s even on PTU. Check APIM retry/spillover pattern (Check 4.1) as mitigation.


💰 CATEGORY 3 — Cost & Efficiency

Check 3.1 — Rate limits configured per deployment (not max)

az cognitiveservices account deployment list \
  --name  --resource-group  \
  --query "[].{name:name, model:properties.model.name, sku:sku.name, capacity:sku.capacity}" \
  -o table

| Finding | Severity | Score | |---------|----------|-------| | Each deployment has explicit TPM capacity set (not maximum) | ✅ Pass | 5 pts | | Single deployment consuming all available quota | ⚠️ Risk | 2 pts | | Unable to determine (no deployments) | N/A | 5 pts |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/quota

Check 3.2 — Model selection diversity (not GPT-4o for everything)

az cognitiveservices account deployment list \
  --name  --resource-group  \
  --query "[].{name:name, model:properties.model.name}" \
  -o table

| Finding | Severity | Score | |---------|----------|-------| | Mix of models (e.g., gpt-4o + gpt-4o-mini, or batch deployments) | ✅ Pass | 3 pts | | Only premium models deployed (no cost-efficient alternative) | ⚠️ Info | 1 pt | | Single deployment only | N/A | 3 pts |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models

Check 3.3 — PTU utilization (if applicable)

Only check if PTU/Provisioned deployments exist:

az cognitiveservices account deployment list \
  --name  --resource-group  \
  --query "[?sku.name=='ProvisionedManaged'].{name:name, capacity:sku.capacity}" \
  -o table

If PTU deployments exist, check utilization via metrics:

az rest --method get \
  --url "https://management.azure.com/providers/microsoft.insights/metrics?api-version=2019-07-01&metricnames=ProvisionedManagedUtilizationV2&timespan=P7D&interval=PT1H&aggregation=Average"

| Finding | Severity | Score | |---------|----------|-------| | PTU utilization 70–85% average (well-sized) | ✅ Pass | 7 pts | | PTU utilization 95% (under-provisioned — users getting 429s) | ⚠️ Risk | 4 pts | | No PTU deployments (using Standard — fine for most startups) | N/A | 7 pts |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/provisioned-throughput

Check 3.4 — Quota utilization (TPM used vs limit)

az cognitiveservices account list-usage \
  --name  --resource-group  \
  -o json

Alternative if the above returns empty (common for newer deployments):

az rest --method get \
  --url "https://management.azure.com/subscriptions//providers/Microsoft.CognitiveServices/locations//usages?api-version=2024-10-01"

| Finding | Severity | Score | |---------|----------|-------| | Quota usage 90% (at risk of rejection — request increase NOW) | ❌ Risk | 0 pts | | Unable to retrieve usage data | ⚠️ Info | 3 pts |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/quotas-limits

> Context from the field: Approved TPM quotas sometimes don't appear in Foundry due to subscription/project mismatches or the new Global/Data Zone quota pooling model. If usage data looks incorrect, verify the quota was assigned to the correct subscription and region.

Check 3.5 — Token consumption trend (tokens per day per deployment)

Requires diagnostic settings sending logs to Log Analytics (Check 2.2). If diagnostics are not configured, skip and note dependency.

AzureDiagnostics
| where TimeGenerated > ago(7d)
| where ResourceProvider == "MICROSOFT.COGNITIVESERVICES"
| where Category == "RequestResponse"
| extend model = tostring(parse_json(properties_s).modelName)
| extend promptTokens = toint(parse_json(properties_s).promptTokens)
| extend completionTokens = toint(parse_json(properties_s).completionTokens)
| summarize 
    DailyPromptTokens = sum(promptTokens),
    DailyCompletionTokens = sum(completionTokens),
    DailyTotalTokens = sum(promptTokens) + sum(completionTokens),
    RequestCount = count()
  by bin(TimeGenerated, 1d), Resource, model
| order by TimeGenerated desc, DailyTotalTokens desc

Use execute_kusto_query with the Log Analytics workspace.

| Finding | Severity | Score | |---------|----------|-------| | Token consumption stable or declining (no runaway growth) | ✅ Pass | 5 pts | | Token consumption growing > 20% day-over-day (investigate) | ⚠️ Attention | 2 pts | | Token consumption spiking > 50% day-over-day (possible prompt leak or loop) | ❌ Risk | 0 pts | | No diagnostic data available (Check 2.2 failed) | ⚠️ Skip | 0 pts — flag dependency |

> 📖 https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/monitoring

> Context from the field: Unmonitored token growth is the #1 surprise cost driver for startups using OpenAI. A single prompt injection loop or chatbot without max_tokens can burn through thousands of dollars in a weekend.


🏗️ CATEGORY 4 — Architecture Patterns

Check 4.1 — AI Gateway / APIM in front of OpenAI (for production)

az apim list --subscription  \
  --query "[].{name:name, sku:sku.name, rg:resourceGroup}" \
  -o table

| Finding | Severity | Score | |---------|----------|-------| | APIM exists with OpenAI-related APIs configured | ✅ Pass | 5 pts | | No APIM — calls go directly to OpenAI endpoint | ⚠️ Risk | 0 pts | | N/A (dev/test environment only) | N/A | 5 pts |

> 📖 https://learn.microsoft.com/en-us/azure/api-management/api-management-authenticate-authorize-azure-openai

Check 4.2 — Environment separation (dev ≠ prod for AI resources)

az cognitiveservices account list \
  --subscription  \
  --query "[?kind=='OpenAI'].{name:name, rg:resourceGroup, tags:tags}" \
  -o json

| Finding | Severity | Score | |---------|----------|-------| | Separate accounts or deployments for dev/prod with tags | ✅ Pass | 5 pts | | Single account with environment tag but shared deployments | ⚠️ Partial | 2 pts | | Single account, no separation, no tags | ❌ Fail | 0 pts |

> 📖 https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ready/landing-zone/design-area/resource-org-subscriptions

Scoring

| Category | Checks | Max Points | |----------|--------|-----------| | 🔐 Security (Identity, Network, Content Filter) | 1.1–1.3 | 35 | | ⚙️ Reliability (Model version, Diagnostics, Locks, Multi-region, 429 monitoring) | 2.1–2.5 | 30 | | 💰 Cost & Efficiency (Rate limits, Model mix, PTU, Quota, Token trend) | 3.1–3.5 | 25 | | 🏗️ Architecture (Gateway, Env separation) | 4.1–4.2 | 10 | | TOTAL | 15 checks | 100 |

Maturity levels

| Score | Level | Meaning | |-------|-------|---------| | 0–39 | 🔴 Critical | Major security and operational gaps — not production-ready | | 40–69 | 🟡 Developing | Some controls in place but significant risks remain | | 70–89 | 🟢 Solid | Good posture — minor improvements recommended | | 90–100 | 🏆 Exemplary | Following WAF best practices for AI workloads |

Expected output

Report header (mandatory — use this exact format)

AI Foundry & OpenAI Posture Check Report

| Field | Value | |-------|-------| | Score | XX / 100 | | Level | 🟡 Developing | | Accounts Assessed | N (list names) | | Deployments Found | M (model details) | | Assessment Date | YYYY-MM-DD |

Per-account findings

For each OpenAI/AI Services account:

| Check | Status | Finding | Recommendation | |-------|--------|---------|----------------| | Managed Identity | ❌ | API keys enabled, no MI | Enable System MI + set disableLocalAuth=true | | Network | ⚠️ | Public access enabled, no firewall | Set defaultAction=Deny + add VNet rules | | ... | ... | ... | ... |

Critical findings (fix immediately)

Security issues that expose the AI endpoint or violate responsible AI principles.

Cost optimization opportunities

Token waste, PTU over-provisioning, or model selection improvements with estimated savings.

Architecture recommendations

Patterns to improve resilience: AI Gateway, retry policies, multi-region deployment.

Remediation guidance

For each ❌ or ⚠️ finding, include in the output:

  1. The specific az CLI command to remediate (suggest only — do not execute)
  2. Use `

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.