A cache hit is not the goal. For a recurring LLM task, the goal is to finish correctly at lower cost or lower elapsed time without weakening access controls. Consider a support assistant that answers questions against a versioned, public product manual. Every request needs the same answer policy and manual, but a different question and a fresh account lookup. Prompt caching can reuse processing for the stable prefix; it does not remember the account, reuse the answer, or certify that the answer is correct. OpenAI describes reuse of computed key-value states for an exact prompt prefix, while new input and output still require work. OpenAI prompt-caching guide
Draw the boundary before sending a request
Before restructuring, this assistant might serialize each request as [timestamp + account ID + latest account lookup] → [answer policy] → [manual v17] → [tool definitions] → [question]. Two questions have different opening tokens, so the expensive shared material no longer forms a common prefix. A better layout for both requests is:
Request 1: [answer policy v3 | stable tool schemas | public manual v17]
<cache boundary> [authorized account facts A, retrieved now | question: return window?]
Request 2: [answer policy v3 | stable tool schemas | public manual v17]
<cache boundary> [authorized account facts B, retrieved now | question: replacement fee?]
The bracketed prefix must be serialized identically: no timestamps, random IDs, request-specific examples, nondeterministic tool order, or silent manual edits. Give the policy and manual deliberate version identifiers; change the bytes and the version together. Put changing retrieval, user text, and the account authorization result after the boundary. Never place account facts in the public shared prefix just because they recur for one user. A cache boundary is an application design point, not portable API syntax: Anthropic supports explicit cache_control on a content block and automatic caching at the last cacheable block; OpenAI's breakpoint options depend on model family. Anthropic prompt-caching guide; OpenAI prompt-caching guide
Do not lengthen a small policy with useless prose merely to clear a threshold. Check the chosen model and request settings first. OpenAI documents a 1,024-visible-token minimum for GPT-5.6 and later but says earlier models vary by settings; Gemini's documented minimums vary by model. Below eligibility, the correct answer is often to leave the short prompt alone. OpenAI prompt-caching guide; Gemini context-caching guide
Price the two answers, including the first write
Here is an invented two-request exercise, not a provider quote or benchmark. Each request has a 2,400-token stable prefix, 300 changing input tokens and 400 output tokens. Hypothetical USD rates per million tokens are $2 uncached input, $2.50 cache write, $0.20 cache read and $8 output. For the uncached pair: 5,400 uncached input + 0 writes + 0 reads + 800 output costs $0.01720. With an eligible write followed by a read: 600 uncached input + 2,400 write + 2,400 read + 800 output costs $0.01408. These buckets are mutually exclusive charges for input tokens, not a write surcharge on top of the same tokens. If both answers pass the same review, that is $0.00860 versus $0.00704 per successful task.
Invented elapsed-time components tell a narrower story. Give each request 100 ms fixed overhead, 400 ms output generation, and 500 ms account-tool time. Assume input processing takes 900 ms on a write or miss and 180 ms on a read. The uncached pair takes 3,800 ms total, or 1,900 ms per successful task; write-then-read takes 3,080 ms, or 1,540 ms per successful task. This is a model of two sequential request outcomes, not a measured service-latency prediction. Output and tool time do not disappear. These figures can be reproduced with a small offline calculator; they are illustrative assumptions, not observed cache behavior.
To reproduce the accounting without API credentials, save the following as synthetic-cache-check.py and run python synthetic-cache-check.py. Its threshold, TTL, rates and times are deliberately invented; its cache key is a teaching simplification, not a provider implementation. A real provider may miss for routing or eligibility reasons even when this toy simulator predicts a hit.
"""Offline illustration only: no SDK, API or provider pricing claims."""
from dataclasses import dataclass
# Entirely invented USD per million tokens and milliseconds per request component.
RATES = {'uncached': 2.0, 'write': 2.5, 'read': 0.2, 'output': 8.0}
MIN_PREFIX = 1000 # Illustrative eligibility threshold, NOT a provider rule.
TTL = 300 # Illustrative five-minute policy, NOT a universal TTL.
@dataclass(frozen=True)
class Request:
tenant: str
version: str
at: int
prefix: int
suffix: int = 300
output: int = 400
success: bool = True
authorized: bool = True
quality: bool = True
def simulate(requests, caching):
cache = {}
records = []
for r in requests:
# Authorization runs independently of cache lookup. Reject before model call.
if not r.authorized:
records.append((r, 'denied', 0, 0, 0, 0, 0))
continue
key = (r.tenant, r.version) # Model/region/config also belong in a production key.
hit = caching and r.prefix >= MIN_PREFIX and key in cache and r.at - cache[key] < TTL
write = caching and not hit and r.prefix >= MIN_PREFIX
if hit or write:
cache[key] = r.at
u = r.suffix + (0 if hit or write else r.prefix)
w = r.prefix if write else 0
read = r.prefix if hit else 0
# Fixed overhead + input processing + output generation + downstream tool time.
latency = 100 + (180 if hit else 900) + 400 + 500
records.append((r, 'hit' if hit else 'write' if write else 'miss', u, w, read, r.output, latency))
return records
def report(name, rows):
good = sum(r.success and r.quality and r.authorized for r, *_ in rows)
u, w, rd, o = (sum(row[i] for row in rows) for i in (2, 3, 4, 5))
dollars = (u*RATES['uncached'] + w*RATES['write'] + rd*RATES['read'] + o*RATES['output']) / 1_000_000
ms = sum(row[6] for row in rows)
print(f'{name}: statuses={[row[1] for row in rows]}, uncached={u}, write={w}, read={rd}, output={o}, cost=${dollars:.5f}, successful={good}, cost/success=${dollars/good:.5f}, total_ms={ms}, ms/success={ms/good:.0f}')
return dollars, ms, good
if __name__ == '__main__':
pair = [Request('A', 'policy-v1', 0, 2400), Request('A', 'policy-v1', 60, 2400)]
before = report('before', simulate(pair, False))
after = report('after', simulate(pair, True))
assert after[0] < before[0] and after[1] < before[1] and after[2] == before[2] == 2
probes = [
('version mismatch', [pair[0], Request('A', 'policy-v2', 60, 2400)], ['write', 'write']),
('below minimum', [Request('A', 'short', 0, 800), Request('A', 'short', 60, 800)], ['miss', 'miss']),
('expired', [pair[0], Request('A', 'policy-v1', 301, 2400)], ['write', 'write']),
('cross tenant', [pair[0], Request('B', 'policy-v1', 60, 2400)], ['write', 'write']),
('cold miss then reuse', pair, ['write', 'hit']),
('unauthorized', [pair[0], Request('A', 'policy-v1', 60, 2400, authorized=False)], ['write', 'denied']),
]
for label, requests, expected in probes:
rows = simulate(requests, True)
actual = [row[1] for row in rows]
assert actual == expected, (label, actual)
print(f'{label}: {actual}')
# Quality flags are external evaluations, not outputs from a synthetic cache.
quality_fail = simulate([pair[0], Request('A', 'policy-v1', 60, 2400, quality=False)], True)
assert quality_fail[1][1] == 'hit' and sum(r.success and r.quality for r, *_ in quality_fail) == 1
print('quality regression: hit, but only 1 of 2 passes; cache hit is not task success')
When a hit is the wrong success metric
For a real pilot, pair comparable questions under the same model, tool policy, manual version and authorization rules. Record provider-reported uncached, write, read and output token buckets where exposed; price each against the applicable model and retention mode. Also record end-to-end elapsed time from request start through tool calls and final accepted answer, time to first token if useful, and whether a reviewer accepted the answer. Divide total billed cost and total elapsed time (including failed attempts and retries) by accepted tasks, not by requests or cache hits. Analyze by route, prefix version, tenant and inter-request gap. A cache hit on an incorrect answer contributes cost and time but no successful task.
Test six paths before rollout: a cold miss followed by reuse; a policy/manual version mismatch; a prefix shorter than the selected model's minimum; a second request after the configured expiry; identical public material arriving under another tenant; and an unauthorized account lookup. The expected cache result for the first request, changed version, short prefix and expired entry is no read. Separate tenants even if the provider's organization-level cache might otherwise permit reuse; authorize before constructing account context, never infer permissions from a read. Repeat answer-quality checks on hits and misses: citations to the manual, correct refusal on insufficient account evidence, and no cross-account facts. The synthetic harness exercises these paths, but cannot establish that a provider behaves this way or that generated answers retain quality.
Provider behavior and retention are design inputs
OpenAI's current guide distinguishes newer explicit/implicit breakpoints and charged writes from earlier implicit-only behavior; its cache lifetime, minimum length and reported usage differ by model. Cached states are machine-local, routing and expiry can cause misses, and caches are not shared across organizations or regional processing boundaries. Its data-retention options also depend on model and organization settings; inspect the current policy rather than assuming an in-memory cache has no retention implications. OpenAI prompt-caching guide
Anthropic exposes cache write and read usage separately and offers five-minute and higher-priced one-hour lifetimes; cache scope is organization- and, on some platforms, workspace-isolated. Its guide describes in-memory key-value representations rather than stored token text, while warning that deletion after a minimum lifetime is not necessarily immediate. Verify its separate data-retention terms for the account. Gemini's Interactions API documents default implicit caching on supported models and usage.total_cached_tokens; explicit cache objects require a different API. Neither a familiar provider name nor an advertised hit rate substitutes for model-specific limits and current pricing. Anthropic prompt-caching guide; Gemini context-caching guide
Ship the prefix change only if matched, reviewed traffic improves cost or latency per accepted answer while refusal, accuracy and account isolation remain intact. If the manual changes frequently, traffic is too sparse for reuse, or the 500 ms account tool dominates the response, optimize that bottleneck instead. The same end-to-end accounting matters in voice AI pipelines, where model input is only one part of the time a user waits.
