LLM MeasurementGitHub

Why did our LiteLLM usage jump, and are the spend logs complete?

Compare two periods of LiteLLM spend logs to see which models and keys drove a change, and which records have incomplete usage.

In this synthetic example, recorded tokens double from 8,000 to 16,000, two keys account for the increase, and the report identifies zero-only and failed records to review.

fleetdiff v0.6.0 reads request-level CSV, JSON or JSONL locally and separates model-by-model usage changes from contributor rankings and completeness.

Run one spend comparison

Requires Git, curl, gh, shasum and tar. Linux or macOS, amd64 or arm64. The sample uses the released binary and a synthetic CSV.

git clone --depth 1 --branch v0.6.0 https://github.com/llm-measurement/fleetdiff.git fleetdiff-litellm-spend
cd fleetdiff-litellm-spend
sh scripts/install.sh v0.6.0 ./release
./release/fleetdiff investigate --litellm-spend examples/litellm-spend/synthetic.csv \
  --before-period 2026-10-07 --after-period 2026-10-08

The headline identifies the model behind the increase:

Recorded tokens doubled (8,000 -> 16,000).
model-1 accounts for 100% of the net recorded increase (+8,000 tokens).
Logged model requests: 8 -> 12.

Who drove it?

The contributor section uses local aliases for the keys:

Who drove the increase?
  key-1: 5,000 -> 10,000 tokens; change +5,000.
  key-2: 1,000 -> 4,000 tokens; change +3,000.
  key-3: 2,000 -> 2,000 tokens; change +0.
  2 leading tracked keys account for 100% of the net recorded increase.
  Attributed tokens: 8,000 -> 16,000.

Are the logs complete?

The completeness section points to records worth checking in LiteLLM:

How complete is the export?
  Records needing usage review (categories may overlap): 2 -> 4.
  Zero-only records, origin unknown: 2 -> 4.
  Failed records (overlap usage categories): 1 -> 2.

To check usage and identity fields in your telemetry, see Do my LLM traces record token usage and user IDs?

Use your own spend logs

Follow the export, install and investigate path to export request-level rows with the read-only SQL projection. Export 16 UTC days for the default comparison of two complete seven-day periods, or choose explicit periods as above.

Choose --group-by team, user, end-user or session for another contributor view. The spend export guide covers the fields and period choices. For continuous metrics and summary windows, use the collector's LiteLLM recipe.

Source and reproduction: released synthetic CSV and accounting and privacy checks. Verification record.

This synthetic example compares logged model requests; zero-only records retain unknown origin, and failed records can overlap usage categories.
Measurement notes