How do I count distinct LLM users without storing user IDs?
Hash each user ID locally, then add the keyed value to a distinct-count sketch. llm-sketchkit stores the measurement state rather than the original IDs, and compatible workers can merge their sketches before estimating the total.
Count a small example
python3 -m venv .venv
.venv/bin/python -m pip install --no-cache-dir llm-sketchkit==0.2.2
.venv/bin/python - <<'PY'
import secrets
from llm_sketchkit import USER_V1, canonicalize_text_v1, hash64, hllpp
from llm_sketchkit.hash import Secret
secret = Secret(secrets.token_bytes(32))
users = hllpp.new("small", USER_V1)
for user_id in ("sample-user-a", "sample-user-b", "sample-user-a", "sample-user-c"):
users.add_hash(hash64(secret, USER_V1, canonicalize_text_v1(user_id)))
print(f"Estimated distinct users: {users.estimate():.0f}")
PY
Estimated distinct users: 3
Across workers and languages
Give compatible workers the same protected secret, profile and hash domain. Merge the sketch states, then estimate once; adding workers' distinct-count estimates would count overlapping users more than once. The Go-to-Python example shows how measurements move into a notebook.
If your application already emits GenAI spans, the collector's user_key mapping provides the collection path. To rank token-heavy users instead, compare user rankings.
Distinct counts are statistical estimates. Your application hashes raw IDs before they reach the sketch.
Measurement notes