LLM MeasurementGitHub

How do I see LLM usage per user without a metric label for every user?

Keep user IDs out of your metric labels. Keep aggregate metrics for totals and rates, a bounded top-users ranking for the biggest contributors, and a distinct-user count whose labels contain no user identities.

Why per-user labels grow so quickly

Metrics backends store each observed combination of label, tag or dimension values as its own series, and hosted services often bill for those combinations. For one metric, series ≈ users × other label combinations when each user appears across those combinations. Model, route, status and instance labels multiply the count further. The exact number depends on which combinations occur; histogram buckets and multiple metrics add more series.

Keep the answers without identity labels

Use aggregate counters for token usage and model attempts, with a small, capped set of model or route labels. Keep the top token-consuming users in a fixed-capacity ranking outside metric labels. Track the number of distinct users as a scalar estimate, also without a user label. This preserves both questions: how many users were active, and which users accounted for most tokens.

For OpenTelemetry users

The released v0.3.1 collector distribution exports counters and distinct-user metrics through Prometheus or OTLP (gRPC or HTTP), so you can use any backend accepting OTLP metrics while keeping keyed rankings in snapshot logs and summary files.

With the released genaisketch connector, map your span's user attribute and enable token-weighted user rankings:

connectors:
  genaisketch:
    fields:
      user_key:
        from_attributes: [enduser.id, user.id]
        canonicalization: text_v1
        domain: user:v1
    topk: 20
    topk_keys:
      - field: user_key
        weight: tokens

This fragment goes into a complete collector configuration with the hashing secret and trace pipeline set up. The distinct-user estimate is a metric value; keyed top-user rankings go to snapshot logs and optional summary files. Keep user identity attributes out of slices. topk limits displayed candidates; the configured sketch profile sets the ranking's fixed capacity.

See the v0.3.1 ranking guide and complete configuration reference. Then use fleetdiff to compare users across windows.

Then check your collector configuration

The secondary check below finds an enduser.id label in a synthetic collector configuration. fleetdiff points to the YAML setting that needs changing.

Requires Git, curl, gh, shasum and tar. Linux or macOS, amd64 or arm64. The installer verifies the v0.5.0 release before extraction.

git clone --depth 1 --branch v0.5.0 https://github.com/llm-measurement/fleetdiff.git fleetdiff-cardinality
cd fleetdiff-cardinality
sh scripts/install.sh v0.5.0 ./release
./release/fleetdiff diagnose - < examples/diagnose/unsafe.yaml

The relevant output is:

2 problems to fix:
  connectors.genaisketch.operation_filter.llm_operations (line 10): Keep only model operations (chat, generate_content, text_completion, embeddings), each once. Tool and agent operations count as model attempts if included.
  connectors.genaisketch.slices[0].keys (line 12): enduser.id is a metric label; remove it from slices and use a hashed field instead.

Exit code 3 means the configuration has a blocking finding. It is the expected result for this example.

Source and reproduction: released synthetic diagnosis fixtures. Verification record.

diagnose reads your collector configuration; your metrics backend reports live series counts.
Measurement notes