How do I detect anomalies in LLM token consumption?
Run fleetdiff scan over retained summary windows to find unusual token use, concentrated users or sessions, tool errors and gaps in usage reporting. Each finding includes the observed value and its baseline.
The synthetic history below produces two unusual windows. Tokens per attempt rise to 3.1 times the baseline; the following window has missing usage on 40% of attempts.
Scan a history
git clone --depth 1 --branch v0.5.0 https://github.com/llm-measurement/fleetdiff.git fleetdiff-anomalies
cd fleetdiff-anomalies
sh scripts/install.sh v0.5.0 ./release
GOWORK=off go run ./examples/scan --out ./synthetic-windows
./release/fleetdiff scan ./synthetic-windows --expected app --recent 2 \
--as-of 2026-10-04T00:30:00Z
Selected output:
2 unusual windows among 2 checked.
2026-10-04T00:28:00Z
reported tokens per attempt: 310.00; typical 100.00 (3.10x).
session-1 now holds 62% of tokens (newly prominent in this scan).
2026-10-04T00:29:00Z
Missing usage: 40.00% of model attempts; typical 0.00% (coverage change, not lower usage).
When a window reports missing usage, check the fields in a local capture: Do my LLM traces record token usage and user IDs?
Exit code 3 means findings were detected; 0 means none were found. The fixed date reproduces this sample. For current traffic, omit --as-of.
Keep watching your application
Retain enough exported windows for a baseline, then run the same command on that directory from a scheduled job. The scan guide explains thresholds, minimum history and persistence. The collector's Kubernetes summary guide covers durable storage and readback.
This example checks planted changes. Review any flag against your own history; sessions that turn over often can look newly prominent.
Measurement notes