Why summaries instead of raw logs
Merge usage across tools, count distinct users, and find the contributors behind a change from retained summaries. In the coding-agent records measured here, 305 days fit in 2.1 MB of summaries, against 1.25 GB of source records. Data and reproduction.
Combine users across tools
Claude Code and Codex contained 46 and 26 user labels, with a union of 52. Adding the counts gives 72. The merged distinct-user estimate stayed within 0.0007 users of the exact union on all 305 days. Method and reproduction.
Query retained history
The summaries answered 793 range queries of 7, 28 and 90 days, with maximum absolute error 0.0034 users. Each query averaged 0.026 ms or less, merging daily states already in memory. Method and reproduction.
Keep state small as identities grow
For 1,000,000 synthetic users, the distinct-count and token-ranking sketches occupied 27,200 to 27,221 bytes, against 26.5 MB of exact keyed-table JSON. Across 3 seeds, observed distinct-count errors were approximately 0.4% to 0.7%, with 100% top-20 recall. Workload and reproduction.
Find concentrated increases
The contributor test flagged days when two leading contributors together accounted for more than half a positive token increase. Compared with the exact analysis, the summary test found 147 matching user flags and 136 matching session flags, with 0 misses and 0 extra flags. It recorded 0 bound violations across 19,862 checked key changes. Rule and reproduction.
Share measurements with an analyst
The serialized-summary scan found 0 hits across 2,076,518 known identifier values in 1,278 files. A separate analyst read summaries with raw-file reads and network access denied by the OS. Its decoded-artifact checks found 0 hits in 855 files, including planted-sentinel checks and injected-leak positive controls. Checks and reproduction.
How it works
These are standard mergeable-sketch properties (Agarwal et al., PODS 2012), also used in Apache DataSketches and BigQuery HLL_COUNT, measured here on LLM usage records.
Data and reproduction
The coding records come from TraceLab v0.0.2, revision 61fea8f97277d0aa247cc2e4c7e31f90ed2bea1c, CC BY 4.0, SyFI Lab, University of Washington. Synthetic scale weights come from successful requests in BurstGPT v2.0, after removing 1,379,449 copied rows; revision 7eb2c4f8350f8a6985272386f5c14af1f678b299, CC BY 4.0, Wang et al. Changes are daily event-time aggregation, keyed summaries, merging and synthetic workloads.
Recorded environment: Apple M4 Max, Darwin 27.0.0 arm64; Go 1.26.9; Python 3.11.14; sketchkit 0.2.2. Collector source: 32bd078e61fefa7294d6d4748a3db4e1aa562994. Replay uses the connector's accounting and event-time clock; exact reference tables validate the answers.
The complete reproduction commands include source and dataset pins. The claim ledger binds each measurement to recorded JSON; run metadata records input hashes, script hashes and machine details.
Daily unions and small-list bytes
Merge the two daily HLL++ states and compare with exact released-label unions. The size comparison is the pair of serialized sketches versus compact JSON for the pair of raw label lists, averaged over days. Use the pinned coding records and recorded environment.
venv/bin/python scripts/data_audit.py
venv/bin/python scripts/features.py
venv/bin/python scripts/sketch_experiments.py
Range queries and retained bytes
Query every contiguous range of 7, 28 and 90 days and compare its union estimate with the exact union. Divide the median of five timed whole batches by each batch's query count. Retain the two platform files per day, excluding the redundant combined view: 2,093,440 summary bytes versus 1,245,532,297 uncompressed source bytes; the source download is 100,939,722 bytes. Use the pinned coding records and recorded environment.
venv/bin/python scripts/sketch_experiments.py
Synthetic scale
Use the pinned hosted-model successful-request weights after removing the copied rows, and seeds 20261010, 20261011 and 20261012. The sampling audit records matched blocks, daily gaps and the weight-sample digest. Each identity occurs once, followed by Zipf(1.2) events. Compare the small HLL++ and frequent-items profiles with an exact keyed table on the same events, in separate processes on the recorded machine.
for seed in 20261010 20261011 20261012; do
./scale data/weights.json "evidence/scale-isolated-1000000-$seed" 1000000 "$seed"
done
Contributor flags and bounds
For every exact key in adjacent daily windows, subtract the before upper bound from the after lower bound and vice versa. Check exact deltas and applicable growth shares against those intervals. Flag when the two largest positive lower-bound changes exceed half of a positive total increase, and compare with the same rule on exact changes. Use the pinned coding records and recorded environment.
venv/bin/python scripts/sketch_experiments.py
Identifier exclusion and analyst isolation
The first scan checks known raw identifiers in serialized summary files. The separate decoded scan checks analyst inputs, outputs and sentinel artifacts, with planted plain and encoded leaks as positive controls. The OS runner verifies raw-file and network denial before running the analyst. Both scan inventories and the denial result appear separately in the recorded evidence. Use the pinned inputs and the macOS recorded environment.
venv/bin/python scripts/isolated_analyst.py
venv/bin/python scripts/privacy_audit.py
venv/bin/python scripts/record_results.py
Try a user and session investigation, a distinct-user count, or the cardinality guide.
Measurement notes
- Compatible identities. Merge measurements with the same identity field, canonicalization, hash domain and key. Start a new comparison period when rotating the key. Report aliases are local to a measurement. The coding records' released user pseudonyms share one namespace across both tools.
- Identity protection. Applications hash IDs before adding them to sketches. Hashes are keyed and pseudonymous, not anonymous: a key holder can test known IDs against exposed hashes, and cardinality sketches can reveal membership (Desfontaines, Lochbihler and Basin, PoPETs 2019). Protect the key, rankings and summary files. Trace fan-out preserves original spans; apply content redaction before forwarding them. The evidence experiment uses a public test key so anyone can reproduce its hashes.
- Estimates and bounds. Distinct counts are statistical estimates; contributor bounds are deterministic. The token-spike example retains every ranked key, giving exact bounds; reports show ranges when bounds differ. The three-user example demonstrates the counting workflow. Measured errors and flag agreement describe the workloads in the evidence bundle.
- Reading findings.
diagnosechecks configuration; live series counts come from your metrics backend.scancompares recent windows, with seasonal patterns and session turnover evaluated against your own history. Flags identify sessions to investigate; confirming a loop takes run-level evidence. Missing usage is a coverage finding, and usage changes identify contributors rather than establish cause. Reported tokens describe usage; they are not an invoice, a cost, or a measure of useful work. - Raw records and summaries. A raw-data system computes exact answers, and raw records remain the source for billing and for questions you didn't choose in advance. Summaries keep selected measurements, which is why they're small. For a few dozen users, raw lists were smaller: 117 versus 103 bytes per day for sketches and lists respectively.
- Scope. The command guides use synthetic inputs with known outcomes. The evidence measurements use one lab's Claude Code and Codex records and synthetic scale workloads on one machine. Sizes are decimal MB and GB; sketch bytes are serialized state, not process memory. Query times are batch averages.