AgentStack
Browse Sign in
Browse Why AgentStack Sell Docs
Sign in
SKILL verified MIT Self-run

Foundry Cost Monitoring

skill-aiappsgbb-awesome-gbb-foundry-cost-monitoring · by aiappsgbb

>

No reviews yet
0 installs
19 views
0.0% view→install

Install

$ agentstack add skill-aiappsgbb-awesome-gbb-foundry-cost-monitoring

✓ scanned · ✓ verified, works with Claude Code, Cursor, and more.

Security review

✓ Passed

No issues found. Passed automated security review. · v0.1.0 How review works →

  • Prompt-injection patterns
  • Secret / credential exfiltration
  • Dangerous shell & filesystem operations
  • Untrusted network calls
  • Known-malicious package signatures

What it can access

  • Network access Used
  • Filesystem access No
  • Shell / process execution No
  • Environment & secrets No
  • Dynamic code execution No

From automated source analysis of v0.1.0. “Used” means the capability is present in the source — more access means more to trust, not that it’s unsafe.

View the full security report →

Verified badge

Passed review? Show it. Paste this badge into your README, it links to the public security report.

AgentStack Verified badge Links to your public security report.
[![AgentStack Verified](https://agentstack.voostack.com/badges/verified.svg)](https://agentstack.voostack.com/security/report/skill-aiappsgbb-awesome-gbb-foundry-cost-monitoring)

Reliability & compatibility

Security review passed
0 installs to date
no reviews yet
1mo ago

Declared compatibility

Claude CodeClaude Desktop

Compatibility is declared by the source manifest. End-to-end runtime verification is coming, see below.

Preview Execution monitoring

We're building live execution health for every listing: tool-call success rate, median latency, uptime, and last-checked timestamps, measured, not self-reported. It isn't live yet, so we don't show numbers we can't stand behind.

How agent discovery & health will work →
Are you the author of Foundry Cost Monitoring? Claim this listing to set pricing, connect Stripe payouts, and keep 70% of every sale.
Sign up to claim

About

Foundry Cost Monitoring

Continuous, FinOps-grade cost monitoring for Microsoft Foundry. Where [foundry-observability](../foundry-observability/SKILL.md) emits the gen_ai.usage.* spans, this skill consumes them — joining them with the Azure Retail Prices API to project per-agent, per-project, and per-tenant cost in near-real time, then wiring Cost Management budgets and Action Groups so the FinOps team sees overruns within minutes instead of next-month's invoice.

> Sibling skills. Pair with foundry-observability for the telemetry > emit path, [paygo-ptu-cost-analyzer](../paygo-ptu-cost-analyzer/SKILL.md) > for one-shot PTU sizing decisions, and > [citadel-spoke-onboarding](../citadel-spoke-onboarding/SKILL.md) when > chargeback is gateway-mediated via APIM x-app-id headers.


§1 — Three layers of cost truth

Foundry cost lands in three places, with three different latencies and three different fidelities:

| Layer | Source | Latency | Fidelity | Use for | |-------|--------|---------|----------|---------| | 1. Foundry portal estimate | ai.azure.com → Operate → Overview / Build → Agent / Model Monitor | ≈ 2–5 min | Estimate — not invoiceable | Near-real-time engineering insight (per-agent rollup) | | 2. OTel real-time projection | App Insights customDimensions['gen_ai.usage.*'] × Retail Prices rate card | ≈ 2–5 min ingest + your KQL window | Engineering projection — not invoiceable | Per-tenant / per-agent / per-conversation projection, anomaly detection, dashboards | | 3. Cost Mgmt actuals | Microsoft.CostManagement/query REST + invoice CSV exports | ≈ 8–24 h | Source of truth for billing | Finance reconciliation, chargeback, budgets |

Reconciliation rule (and this is the rule that prevents the awkward quarterly call with finance): treat layer 3 as authoritative for anything that ends up on the bill. Layers 1 and 2 are engineering tools to see overruns before layer 3 catches up. Per the MS Learn cost guidance: "treat your invoice and meter records as the source of truth."


§2 — Pricing source: Azure Retail Prices REST API

The Retail Prices API is anonymous, public, and rate-card complete:

GET https://prices.azure.com/api/retail/prices
    ?$filter=
    [&$top=N]
    [&$skip=N]

No auth header. No subscription. Just curl (or urllib.request). The filter shape that matters for Foundry models in Sweden Central:

# Foundry models (post-rename — current GA service classification)
curl -s "https://prices.azure.com/api/retail/prices?\$filter=serviceName eq 'Foundry Models' and armRegionName eq 'swedencentral' and contains(meterName, 'Tokens')" | jq '.Items[0:3]'

# Embeddings — Sweden Central requires GlobalStandard SKU; filter accordingly
curl -s "https://prices.azure.com/api/retail/prices?\$filter=serviceName eq 'Foundry Models' and armRegionName eq 'swedencentral' and contains(productName, 'embedding')"

Each Items[] row carries the fields you need to price a token row:

| Field | Example | What it means | |-------|---------|---------------| | serviceName | Foundry Models | Top-level service classification | | productName | Azure OpenAI, Azure Llama Models, Azure OpenAI GPT5 | Model family bucket | | skuName | gpt 4.1 Inp regnl | Deployment-shape SKU | | meterName | gpt 4.1 Inp regnl Tokens | Billing meter | | retailPrice | 0.00242 | Per unitOfMeasure | | unitOfMeasure | 1K, 1M | Token unit (mix of 1K and 1Mnormalize before multiplying) | | currencyCode | USD | Pricing currency | | armRegionName | swedencentral | Region | | type | Consumption | Consumption for PAYGO, Reservation for PTU |

Pagination: response includes NextPageLink; iterate until null when caching the full catalog (~thousands of rows for AI + Machine Learning).

> Service-name drift. The Retail Prices API used to surface OpenAI > meters under serviceName eq 'Cognitive Services'. Current GA > classification is Foundry Models. The drift detector for this > skill polls the Retail Prices endpoint weekly — if a new classification > appears, refresh § 11 of this skill and bump PATCH.

Cache TTL recommendation: 24 h. Rates change weekly at most; daily refresh comfortably beats invoice cadence.


§3 — Foundry project tag chargeback (Preview)

Foundry automatically tags Models-sold-by-Azure usage records with a project tag whose value is the Foundry project name. The chargeback flow is then:

  1. In Cost Management → Cost Analysis, scope to your Foundry resource.
  2. Add filterTagproject.
  3. Pick one or more projects → cost split by project, time-series.

CLI equivalent:

az consumption usage list \
  --subscription "$SUB" \
  --start-date "$(date -u -v-30d +%Y-%m-%d)" \
  --end-date "$(date -u +%Y-%m-%d)" \
  --query "[?tags.project=='proj-fsi-claims'].{date:usageStart, meter:meterDetails.meterName, qty:quantity, cost:pretaxCost}" \
  -o table

> Verbatim preview limit (MS Learn, manage-costs#chargeback-with-project-level-cost-attribution-preview): > > "Project-level cost attribution is currently supported for Models > sold by Azure (Azure Direct models, including Azure OpenAI). It isn't > yet supported for models served through Azure Marketplace."

Concretely: Azure OpenAI / Foundry-native models → tagged automatically. Llama / Mistral / DeepSeek / Cohere via Azure Marketplace → no project tag; fall back to layer-2 OTel projection (§ 4) or per-resource grouping (one project ↔ one resource group).

The tag is not manually applied — do not try to add it via az tag create-or-update; Foundry control plane writes it on the underlying meter records.


§4 — In-flight cost projection: KQL on gen_ai.usage.*

This is the engineering-insight layer. Requires foundry-observability already wired (account-level App Insights connection + ACA-side configure_azure_monitor()). The MAF 1.6.0+ schema emits OTel spans with the following customDimensions:

| Key | Type | Example | |-----|------|---------| | gen_ai.system | string | azure_openai | | gen_ai.request.model | string | gpt-5.4-mini | | gen_ai.response.model | string | gpt-5.4-mini-2026-01-15 | | gen_ai.usage.input_tokens | int | 1284 | | gen_ai.usage.output_tokens | int | 342 | | gen_ai.operation.name | string | chat, embeddings |

Spans land in traces / dependencies depending on instrumentation. The canonical filter for MAF-emitted spans is cloud_RoleName == "agent_framework" (verified at MAF 1.6.0+ — see [foundry-observability](../foundry-observability/SKILL.md) § "OTel cloud_RoleName" for the rationale).

4.1 Per-agent cost rollup (last 24 h)

// Per-agent token cost (last 24h) — projection-only, not invoiceable
let rateCard = externaldata(model:string, in_rate_usd_per_1k:real, out_rate_usd_per_1k:real)
  [@"https://raw.githubusercontent.com///main/rate-card.csv"]
  with (format="csv", ignoreFirstRecord=true);
dependencies
| where timestamp > ago(24h)
| where cloud_RoleName == "agent_framework"
| where isnotempty(customDimensions["gen_ai.usage.input_tokens"])
| extend
    agent      = tostring(customDimensions["gen_ai.agent.name"]),
    model      = tostring(customDimensions["gen_ai.request.model"]),
    in_tokens  = toint(customDimensions["gen_ai.usage.input_tokens"]),
    out_tokens = toint(customDimensions["gen_ai.usage.output_tokens"])
| join kind=leftouter rateCard on model
| extend
    in_cost  = (in_tokens  / 1000.0) * in_rate_usd_per_1k,
    out_cost = (out_tokens / 1000.0) * out_rate_usd_per_1k
| summarize
    calls            = count(),
    input_tokens     = sum(in_tokens),
    output_tokens    = sum(out_tokens),
    projected_cost_usd = sum(in_cost + out_cost)
    by agent, model
| order by projected_cost_usd desc

4.2 Per-project cost rollup

gen_ai.agent.name carries the agent name. To get per-project rollup without the preview tag, instrument the project name as a resource attribute on the OTel exporter (e.g. OTEL_RESOURCE_ATTRIBUTES="foundry.project=proj-fsi-claims") and group on customDimensions["foundry.project"] instead of agent. Where the project tag (§ 3) IS available, prefer Cost Management for authoritative reconciliation and reserve KQL for in-flight projection.

4.3 Per-tenant cost rollup (Citadel APIM gateway path)

If traffic goes through the AI Citadel gateway (see § 5), the APIM fragment-policy stamps the tenant identifier into the OTel span as enduser.id (gateway maps x-app-id header → enduser.id). The per-tenant rollup is identical to § 4.1 but groups on tostring(customDimensions["enduser.id"]).


§5 — Citadel APIM x-app-id chargeback path

When Foundry traffic is fronted by an APIM AI Gateway (the AI Citadel Hub topology), tenant identity arrives on the x-app-id request header and the gateway logs it into App Insights. This is the cleanest chargeback path for multi-tenant SaaS deployments because it does not depend on Foundry's project-tag preview.

Wiring path (full detail in [citadel-spoke-onboarding](../citadel-spoke-onboarding/SKILL.md)):

Tenant request → APIM (validates JWT, stamps x-app-id) → Foundry
                                                            ↓
                                                    OTel span tagged with
                                                    enduser.id = x-app-id

KQL aggregates over enduser.id then rejoin to the customer-ID-to-name table (kept outside this skill — usually a Cosmos/Kusto reference table).


§6 — Budget alerts via Azure Cost Management

Set bottom-up: subscription budget + per-RG budgets for spike isolation.

6.1 Bicep (recommended)

// budget.bicep — per-RG budget with notifications at 80% / 100% / 120%
@description('Budget amount in USD per month')
param budgetAmount int = 5000

@description('FinOps distribution list')
param finopsEmails array = [
  'finops@.com'
]

resource budget 'Microsoft.Consumption/budgets@2024-08-01' = {
  name: 'budget-foundry-${resourceGroup().name}'
  properties: {
    amount: budgetAmount
    timeGrain: 'Monthly'
    timePeriod: {
      startDate: '2026-06-01'
    }
    category: 'Cost'
    notifications: {
      Warning_80: {
        enabled: true
        operator: 'GreaterThan'
        threshold: 80
        contactEmails: finopsEmails
        thresholdType: 'Actual'
      }
      Critical_100: {
        enabled: true
        operator: 'GreaterThan'
        threshold: 100
        contactEmails: finopsEmails
        contactGroups: [
          actionGroup.id
        ]
        thresholdType: 'Actual'
      }
      Forecast_120: {
        enabled: true
        operator: 'GreaterThan'
        threshold: 120
        contactEmails: finopsEmails
        thresholdType: 'Forecasted'
      }
    }
    filter: {
      tags: {
        name: 'project'
        operator: 'In'
        values: [
          'proj-fsi-claims'
          'proj-fsi-onboarding'
        ]
      }
    }
  }
}

resource actionGroup 'Microsoft.Insights/actionGroups@2024-10-01-preview' = {
  name: 'ag-finops-${resourceGroup().name}'
  location: 'global'
  properties: {
    groupShortName: 'finops'
    enabled: true
    webhookReceivers: [
      {
        name: 'finops-teams'
        serviceUri: ''
        useCommonAlertSchema: true
      }
    ]
  }
}

6.2 CLI equivalent

# Resource-group scoped budget — fires Action Group on actual cost > $5000
az consumption budget create-with-rg \
  --resource-group "$RG" \
  --budget-name "budget-foundry-$RG" \
  --amount 5000 \
  --time-grain Monthly \
  --start-date "$(date -u +%Y-%m-01)" \
  --category Cost \
  --notifications-properties \
      "[{operator:'GreaterThan',threshold:80,contactEmails:['finops@.com'],thresholdType:'Actual'}]"

> Hard-limit caveat (MS Learn manage-costs § Create budgets): > "Azure OpenAI doesn't currently provide [hard-limit] functionality. > You can start automation from action groups as part of your budget > notifications to take more advanced actions, but this functionality > requires additional custom development." > > Budgets are alerts, not throttles. For hard ceilings, wire the > Action Group → Logic App / Function that disables the offending > deployment (e.g. az cognitiveservices account deployment update > --properties.callRateLimit=0).


§7 — Action Group → Webhook / Logic App / Teams

Standard Azure pattern. The Action Group from § 6.1 wires three common downstream channels:

| Channel | Use case | |---------|----------| | Teams incoming webhook | FinOps channel notification (useCommonAlertSchema: true so payload is parseable) | | Logic App | Disable / scale-down offending deployments automatically | | Azure Function | Custom remediation (e.g. notify owner, file ticket, page on-call) |

The Action Group is fire-and-forget — Cost Management posts the budget breach event payload to every receiver in parallel.


§8 — Anomaly detection on token cost

The KQL queries in § 4 are point-in-time. To catch unexpected spend spikes (a leaked API key, a runaway agent loop, a prompt that exploded context-window utilisation), wrap the time-series in series_decompose_anomalies:

// 7-day token-cost anomaly detection per agent
let bin_size = 1h;
dependencies
| where timestamp > ago(7d)
| where cloud_RoleName == "agent_framework"
| where isnotempty(customDimensions["gen_ai.usage.input_tokens"])
| extend
    agent      = tostring(customDimensions["gen_ai.agent.name"]),
    model      = tostring(customDimensions["gen_ai.request.model"]),
    in_tokens  = toint(customDimensions["gen_ai.usage.input_tokens"]),
    out_tokens = toint(customDimensions["gen_ai.usage.output_tokens"])
| summarize total_tokens = sum(in_tokens + out_tokens) by agent, bin(timestamp, bin_size)
| make-series series=sum(total_tokens) default=0 on timestamp from ago(7d) to now() step bin_size by agent
| extend (anomalies, score, baseline) = series_decompose_anomalies(series, 2.5)
| mv-expand timestamp, series, anomalies, score, baseline
| where toint(anomalies) != 0
| project agent, timestamp=todatetime(timestamp), tokens=toint(series), anomaly=toint(anomalies), score=todouble(score)
| order by score desc

Wire this query to an Azure Monitor scheduled query alert (15-min frequency, threshold count() > 0) → same Action Group from § 6.


§9 — Reconciliation drift: OTel projection vs Cost Mgmt actuals

The two layers will never match exactly. Expected drift:

| Source | Latency | Causes of drift vs invoice | |--------|---------|----------------------------| | OTel projection | 2–5 min ingest | Stale rate card; dropped spans; OTel sampling; cached input pricing not modeled; reservation discounts not applied | | Cost Mgmt actuals | 8–24 h | Authoritative — the bill |

Acceptable drift: ±5% day-over-day, ±2% month-over-month. If your projection consistently overshoots, the rate card is stale (refresh from Retail Prices API). If it consistently undershoots, you're missing a meter — usually the cached-input meter (some models split cached input into a separate, cheaper meter) or the fine-tuned-hosting hourly meter.

Reconciliation cadence:

  1. Daily: run § 4.1 KQL for previous day, compare to query.usage

result for the same scope + day. Log delta.

  1. Weekly: refresh the rate-card CSV from Retail Prices.
  2. Monthly: full reconciliation against invoice CSV export. Adjust

series_decompose_anomalies baseline if drift bias exceeds ±2%.


§10 — Common pitfalls

| Pitfall | Symptom | Fix | |---------|---------|-----| | Model-name normalization | KQL model join misses (gpt-5.4-mini

Source & license

This open-source skill is cataloged on AgentStack and links to its original source — we do not rehost the code.

Install and usage instructions live in the source repository linked above.

Reviews

No reviews yet, be the first.

Versions

  • v0.1.0 Imported from the upstream source.