LLM MeasurementGitHub

How do I count distinct LLM users without storing user IDs?

Hash each user ID locally, then add the keyed value to a distinct-count sketch. llm-sketchkit stores the measurement state rather than the original IDs, and compatible workers can merge their sketches before estimating the total.

Count a small example

Requires Python 3.11 or later. Installation downloads the pinned release from PyPI; counting runs locally.

python3 -m venv .venv
.venv/bin/python -m pip install --no-cache-dir llm-sketchkit==0.2.2
.venv/bin/python - <<'PY'
import secrets
from llm_sketchkit import USER_V1, canonicalize_text_v1, hash64, hllpp
from llm_sketchkit.hash import Secret

secret = Secret(secrets.token_bytes(32))
users = hllpp.new("small", USER_V1)
for user_id in ("sample-user-a", "sample-user-b", "sample-user-a", "sample-user-c"):
    users.add_hash(hash64(secret, USER_V1, canonicalize_text_v1(user_id)))
print(f"Estimated distinct users: {users.estimate():.0f}")
PY
Estimated distinct users: 3

Across workers and languages

Give compatible workers the same protected secret, profile and hash domain. Merge the sketch states, then estimate once; adding workers' distinct-count estimates would count overlapping users more than once. The Go-to-Python example shows how measurements move into a notebook.

If your application already emits GenAI spans, the collector's user_key mapping provides the collection path. To rank token-heavy users instead, compare user rankings.

Source: llm-sketchkit v0.2.2. The checked command above defines the complete synthetic input. Verification record.

Distinct counts are statistical estimates. Your application hashes raw IDs before they reach the sketch.
Measurement notes