LLM MeasurementGitHub

LLM Measurement

LLM Measurement is a set of open-source tools for understanding LLM usage: what changed, who drove it, and whether the usage records support the answer. They work from the telemetry you already collect, without a metric label for every user, and their outputs carry keyed hashes instead of raw user or session IDs.

They are designed with care and built on peer-reviewed sketching algorithms. Rankings come with guaranteed bounds, distinct counts come with a known statistical error, and missing usage is counted separately as missing. Every result on this site comes with the commands to reproduce it.

Guides

Each guide runs released versions on synthetic examples.

Why summaries instead of raw logs? See the measured sharing and retention properties, with source data and methods.

fleetdiff terminal report showing a flagged session in a synthetic token-usage comparison
fleetdiff's synthetic session example. Run it with the released binary.

Using LiteLLM?

Export request-level spend logs, install fleetdiff v0.6.0, and investigate what changed. Start with the released export, install and investigate path, or try a synthetic spend export.

For continuous measurement, use the collector's LiteLLM recipe to create metrics and summaries as traffic arrives.

Choose a tool

fleetdiff reads local spend exports, trace captures or summary files. The Collector connector creates metrics and summaries continuously. llm-sketchkit supplies the measurement library for Go services and Python analysis.