Which user is driving my LLM token spike?
Usage jumped? Compare two periods and see which users and sessions account for the increase.
In this synthetic example, one user adds 3,000 reported tokens while the application's total rises from 400 to 3,300.
fleetdiff shows both the contributors and the split between more model attempts and more tokens per attempt.
To see who uses the most overall, rather than who changed, see Who is tokenmaxxing?
Run the comparison
git clone --depth 1 --branch v0.5.0 https://github.com/llm-measurement/fleetdiff.git fleetdiff-token-spike
cd fleetdiff-token-spike
sh scripts/install.sh v0.5.0 ./release
./release/fleetdiff investigate --before examples/sessions/data/before \
--after examples/sessions/data/after --expected app
The report opens with:
1 of 8 tracked sessions flagged for review: 90.91% of attributed tokens.
Reported tokens: 400 -> 3300; model attempts: 4 -> 13.
+1592 tokens from attempt count; +1308 tokens from tokens per attempt.
Tokens per attempt: 100.00 -> 253.85.
The user section identifies the leading contributor:
Which tracked users contribute tokens or model attempts?
key / measurement before -> after after share review
user-1 / top_users 0 -> 3000 90.91%
Use your own application
Enable token-weighted user_key and session_key rankings in the released collector, retain two complete summary windows, and pass their directories as --before and --after. Set --expected to your actual producer inventory and keep the hashing key consistent across the comparison.
For a gateway setup, start with the LiteLLM recipe. To keep watching, scan retained history.
A flag marks a session worth a look. With real data, the report shows ranges wherever bounds differ.
Measurement notes