AI Governance

Building an AI Landing Zone on Azure — Part 4: Metering every call

The usage pipeline behind the AI gateway: APIM to Event Hub to a Function App to Cosmos DB, managed identity end to end.

This is part 4 of a five-part series on building an AI Landing Zone on Azure. Part 1: Why every enterprise needs an AI gateway · Part 2: The platform underneath · Part 3: One front door for many models · Part 4: Metering every call (this post) · Part 5: Turning tokens into euros

Metering every call: a pipeline full of glowing particles passing a metering station whose gauge reads zero, above a blank ledger drawer

Parts 1 to 3 described building a private gateway that every AI call goes through. This part is about the thing that makes the gateway worth more than a fancy reverse proxy: for every call, it records exactly what was consumed, by whom, and writes that somewhere a (finance) team can use.

This is also the part of the project that took the most effort relative to how simple it looks on the diagram, so I am going to spend as much time on what went wrong as on what the design is. If you only read one section, read the one about silent failures.

The pipeline

Usage metering pipeline: APIM emits a usage record per call to Event Hub, a Flex Consumption Function upserts it into Cosmos DB, Power BI prices it; Application Insights and VictoriaMetrics as side channels, managed identity on every hop

Hop 1: APIM emits the usage record. In the outbound section of every API policy, a fragment runs after a successful response. It reads the model and usage information, builds a JSON document, and posts it to the Event Hub REST endpoint with a bearer token obtained through APIM's user-assigned managed identity. Only 2xx responses produce a record; a 404 or 429 upstream is not considered as usage.

Hop 2: Event Hub. One hub, ai-usage, with a dedicated consumer group for ingestion, public access disabled, private endpoint in the zone. Retention is a day in non-production and a week in production, which is enough to survive a Function outage without losing records.

Hop 3: the Function App. A Python Function App on the Flex Consumption plan, VNet-integrated into its own subnet, with a user-assigned identity. It has an Event Hub trigger that batches records and upserts them into Cosmos DB, and pushes the same numbers as Prometheus metrics to VictoriaMetrics on the AKS cluster so the platform team has a live view.

Hop 4: Cosmos DB. A SQL API account, public access disabled, local authentication disabled. The usage container is partitioned by product name, because product is what every report groups by.

Hop 5: the streaming side channel. For routes that stream and are not buffered, there is no usage block for APIM to read, so a second function on a five-minute timer queries a token metric from Application Insights and writes aggregate records.

The usage record

The record APIM emits is deliberately richer than "tokens in, tokens out", because everything you want to report on later has to be captured here. It carries the APIM subscription and product (for chargeback), the model as returned by the API and the deployment name as requested, the gateway name and region, the operation and backend, and then the consumption itself: prompt tokens, cached input tokens, cache-write tokens, completion tokens, reasoning tokens and total tokens for language models, or pages processed and document size for OCR. It also carries a usageSource field that says whether the numbers came from the response body (exact) or from an estimate, and an isStream flag.

Two details here saved me later. Capturing cached and cache-write tokens separately mattered as soon as a model appeared (GTP 5.6) with a separate price for cache writes. And the fragment reads both the Chat Completions shape (prompt_tokens, completion_tokens) and the Responses API shape (input_tokens, output_tokens, with nested details), so the same fragment meters both the document-extraction workload and the coding assistant.

Exact versus estimate: the streaming trade-off

In part 3 I mentioned that whether APIM buffers a response is a per-route decision. What makes that important?

When a chat completion is not streamed, the response is one JSON body with a usage block. APIM reads it in the outbound policy and the record is exact.

When a chat completion is streamed, the response is a series of server-sent events, and unless the backend has been asked to include usage in the final chunk and APIM has buffered the whole stream, there is nothing for the outbound policy to read. You can buffer, in which case the fragment parses the terminal event and you still get exact numbers, at the cost of the caller waiting for the full answer. Or you do not buffer, in which case the fragment falls back to the estimate that APIM's own token-limit policy computed on the way in, marks the record usageSource = estimate, and the caller gets live token-by-token output.

I made this a per-route toggle in the policy rather than a global choice. The coding-assistant route buffers, because those calls are large. Interactive routes stream. As a safety net, the OpenAI policy also emits a token metric to Application Insights on every call, and the timer function scrapes that metric every five minutes and writes aggregate "streaming" records, so even routes that stream without buffering produce a total-tokens number. Those streaming records have no input/output split, which has a consequence for pricing that I will cover in part 5.

The three silent failures

Now the part that took me the most time to figure out. Each of these produced the same symptom: every component reported healthy, every test call returned 200, and the Cosmos container stayed empty. None of them threw an error I could see.

1. The Event Hub logger that does nothing

APIM has a native log-to-eventhub policy backed by a logger resource, and that logger supports managed identity authentication. It is the obvious way to send events to Event Hub. I configured it, the logger resource created fine, the policy validated, calls returned 200, and nothing ever arrived in the hub.

Against a private Event Hub, with public access disabled, the managed-identity logger silently no-ops. I do not know whether that is a limitation of how the logger resolves the endpoint or something about how it acquires the token, and Microsoft's documentation is not clear on it. What I do know is that after adding diagnostic response headers to surface internal policy state, the logger path was being reached and simply had no effect.

The fix was to stop using the logger and emit the record with a send-request policy instead: acquire a token for the Event Hubs audience with authentication-managed-identity, POST the JSON to the hub's REST endpoint. That works reliably, and it has a nice feature that the logger does not have: the endpoint is a named value, so the same fragment works in every environment. Both the token acquisition and the send have ignore-error set so that a metering hiccup never fails a caller's request, and I left a note in the fragment that flipping those to false is the fastest way to see the real error when debugging.

2. The empty setting that on the Function App

The Function App uses managed identity for its host storage as well, because shared keys are disabled on the storage account. That is configured with a triple of settings: AzureWebJobsStorage__accountName, AzureWebJobsStorage__credential = managedidentity, and AzureWebJobsStorage__clientId.

At one point, trying to stop the platform from re-injecting a connection-string value, I set the bare AzureWebJobsStorage key to an empty string. The Function App started. The Event Hub trigger attached. Invocation counts stayed at zero. Event Hub metrics showed Incoming Messages climbing steadily and Outgoing Messages flat at zero.

When AzureWebJobsStorage exists, even as an empty value, the Functions host uses it and ignores the three AzureWebJobsStorage__* settings that configure managed-identity access. The host cannot checkpoint, so the trigger never consumes, and because the host considers itself healthy there is no error to find. The fix is to remove the empty setting entirely, not set it to anything. This one is now the first line in the operations handover, because the symptom (incoming up, outgoing zero) is so specific that it should be the first thing anyone checks.

3. Contributor is not a data reader

Cosmos DB has two separate permission systems. Azure RBAC (Owner, Contributor, Reader) controls the control plane: the account, the containers, the throughput. The data plane, actually reading and writing documents, is governed by Cosmos SQL role definitions, assigned with a separate command, and Azure RBAC is not used.

"But I'm Contributor, why can't I read"? The permission errors are easy to misread as a networking problem, because everything else that fails does so for networking reasons. The fix is the built-in Cosmos DB Built-in Data Contributor role for the Function App identity and Data Reader for the reporting identity, assigned at the data plane, which Terraform can do, and which I added to the IaC code obviously.

Two smaller things that also cost time

The Flex Consumption plan does not support WEBSITE_RUN_FROM_PACKAGE, so the code is deployed with a zip deploy triggered from Terraform. A terraform_data resource hashes the function source, and when the hash changes, a provisioner exports pinned requirements with uv, downloads copies of every Python dependency for the Function App's runtime and bundles them with the code, zips the result and pushes it. It is not elegant, but it means a change to the ingestion code goes through the same merge request and pipeline as everything else.

It looks something like this:

resource "terraform_data" "function_build_deploy" {
  triggers_replace = [
    local.fn_src_hash,
    azurerm_function_app_flex_consumption.ingestion.id,
  ]

  provisioner "local-exec" {
    interpreter = ["/bin/bash", "-c"]
    working_dir = local.fn_src_dir
    command     = <<-EOT
      set -euo pipefail

      # 1. Pin deps from the uv lock.
      uv export --no-hashes --no-dev --format requirements-txt -o requirements.txt

      # 2. Vendor packages for the target runtime.
      rm -rf .python_packages
      uv pip install \
        --target .python_packages/lib/site-packages \
        --python-version 3.14 --only-binary=:all: \
        -r requirements.txt

      # 3. Build the package, excluding local-only stuff.
      mkdir -p ../.terraform.tmp
      rm -f ../.terraform.tmp/function_pkg.zip
      zip -r -q ../.terraform.tmp/function_pkg.zip . \
        -x '.venv/*' -x '__pycache__/*' -x '*.pyc' -x '.python_packages/**/__pycache__/*'

      # 4. Deploy without remote build.
      az functionapp deployment source config-zip \
        --resource-group "${azurerm_resource_group.ailz.name}" \
        --name "${azurerm_function_app_flex_consumption.ingestion.name}" \
        --src ../.terraform.tmp/function_pkg.zip \
        --build-remote true
    EOT
  }
}

And APIM policy expressions are C# with a type-locked variable store. The built-in token-limit policy writes its token estimate as a 64-bit integer; an early version of my fragment read it as a 32-bit one, and the collision produced a policy error that only appeared on some routes. A conversion in the fragment solved it, but it is a reminder that policy expressions are real code and deserve the same care.

The runbook that came out of it

Everything above went into a troubleshooting section at the top of the operations handover, each step tells you which side of the pipeline is broken:

  1. Are records actually missing, or is the report stale? Count recent documents in Cosmos.
  2. Is the gateway returning 200? Only successful calls are metered.
  3. Is the hub flowing? Incoming flat means the policy is not emitting; incoming up and outgoing zero means the Function is not consuming.
  4. Data present but report stale? Refresh, and check the rate card is seeded.

Next: turning it into Euros

The pipeline now produces one trustworthy record per call, with enough detail to price it. Part 5 is about the rate card, computing cost at query time, handling cached and reasoning tokens, and the report that finally gave the finance team a monthly AI bill per cost centre.

If you have hit any of these three failures yourself, I would enjoy hearing about it; and if you would rather not hit them at all, this pipeline is part of the landing zone accelerator I am packaging. Connect with me on LinkedIn or reach out via ascode.nl.

Written by

Erik Christiaans

Independent cloud and AI platform architect in Velsen, NL. Twenty years of Azure, AWS, Entra ID and Kubernetes for insurers, banks and other regulated enterprises - and the control mappings that make those platforms defensible.

Book a 30-minute call