LLM MeasurementGitHub
Guides

Prompt cache hit rate dropped after an update? Find whose uncached input grew

Compare recorded cache reuse before and after a change, and find which keys or sessions added uncached input.

What you'll see

Example output from the synthetic sample below; your values will reflect your input.

Recorded cache-read share falls from 92% to 61%, adding 620 uncached input tokens.

But cache reuse fell: cache-read share 92% -> 61% (uncached input +620).

Which key's uncached input grew?

The keyed uncached-growth ranking uses requests with usable cache-read detail and a group identity. In this fixture, key-1 rises from 80 to 700 uncached input tokens while the other key stays flat.

  key-1: 80 -> 700 uncached input tokens; change +620.

The report's local aliases refer to the same identities in its total-token and cache rankings. Rankings carry bounds when the sketches cannot identify an exact change.

How much input was reused?

Both periods record 2,000 input tokens and 200 output tokens. Every request reports cache-read and cache-write detail, so the denominators cover all input in this example:

    Cache-read share of recorded input: 92% -> 61%.
      Cache reads: 1,840 of 2,000 input tokens -> 1,220 of 2,000 (all requests report cache detail).
    Cache-write share of recorded input: 5% -> 7.5%.
      Cache writes: 100 of 2,000 input tokens -> 150 of 2,000 (all requests report cache detail).

Uncached input is input minus cache reads; it includes cache writes. Cache reads and writes are subsets of recorded input. Shares measure input tokens.

Check coverage before interpreting a change

Each cache field has its own covered-input denominator and request count. Missing, invalid and subset-violating details are excluded and counted separately; requests without group identities stay outside the ranking. Explicit zero detail is recorded zero, not proof of provider origin. A missing detail or zero denominator leaves the share unavailable.

If the recorded share falls, review the provider's cache diagnostics and whether the gateway route preserves cache markers.

No data handy? Try the sample

Requires Git, curl, gh, shasum and tar. Linux or macOS, amd64 or arm64. The sample uses the released fleetdiff v0.7.0 binary and a four-request synthetic CSV. No Go, collector, model key or paid call is needed.

git clone --depth 1 --branch v0.7.0 https://github.com/llm-measurement/fleetdiff.git fleetdiff-cache-check
cd fleetdiff-cache-check
sh scripts/install.sh v0.7.0 ./release
./release/fleetdiff investigate --litellm-spend examples/litellm-spend/cache-drop.csv \
  --before-period 2026-10-07 --after-period 2026-10-08 --group-by key

Source and reproduction: released synthetic cache-drop CSV and accounting and privacy checks. Verification record.

Uncached-input growth can reflect input volume or workload mix, not necessarily a per-key cache-share loss. These recorded shares do not diagnose causes or costs; the synthetic example checks arithmetic, not provider cache behavior. Collector summaries support aggregate shares with sufficient coverage, but lack per-model cache counters and cache-weighted key/session rankings.
Measurement notes. Cache measurements and limits.