- Mark the stable tools + system prefix with explicit
cache_controlso repeats stop billing as full input. - Clear your model’s minimum token floor (512 on Opus 5 / Fable 5 / Mythos 5 class) or the cache never writes.
- Verify with usage fields: first call creates cache tokens; second call within TTL reads them.
Static system prompts and tool schemas keep getting billed as full input because nothing marks where the reusable prefix ends. Prompt caching only helps when you place an explicit cache_control breakpoint on that stable prefix, and when the prefix clears your model’s minimum token floor.
Why it matters
If your agent loop ships the same system doc and tool schemas on every turn, you are paying base input for a prefix that should be a cache read. The breakpoint is the control surface: mark the end of the stable prefix, keep the user turn uncached, and confirm a write then a read in usage. The silent failure mode is the trap. Under-minimum prefixes never throw; they just never cache. Check the minimum table when you change model ids, because Opus 5 at 512 is not the same floor as Sonnet 5 at 1,024 or Haiku 4.5 at 4,096. Before you trust savings, run two identical calls within the TTL and prove cache_read_input_tokens moves.
Cache reads bill at a fraction of base input (typically 0.1x; check current pricing for your model). Model minima differ: Claude Opus 5, Claude Fable 5, and Claude Mythos 5 (plus 5.1 variants) need at least 512 tokens; Claude Sonnet 5 and Claude Sonnet 4.6 need 1,024; Claude Opus 4.6 / 4.5 and Claude Haiku 4.5 need 4,096. Shorter prefixes marked with cache_control process with no error, and both cache_creation_input_tokens and cache_read_input_tokens stay 0. Place breakpoints on the last block that stays identical across calls, not on per-request timestamps or the live user turn.
How to set prompt-cache breakpoints
Use the Messages API host https://api.anthropic.com/v1/messages, model claude-opus-5 (512-token minimum), and put cache_control: {"type": "ephemeral"} on the last static system block and on the last tool definition. Leave the changing user message uncached. Do not paste a placeholder API key.
- Export
ANTHROPIC_API_KEYand fail closed if it is unset. - Build a large, stable
SYSTEM_DOC(instructions plus real static docs) that clears 512 tokens forclaude-opus-5. - Attach
cache_controlon the last tool and the system text block; leave the user turn uncached. - Send the request twice within the TTL. Expect
cache_creation_input_tokensgreater than 0 on the first call andcache_read_input_tokensgreater than 0 on the second. If both stay 0, lengthen the static system text or re-check the model minimum.
#!/usr/bin/env bash
set -euo pipefail
: "${ANTHROPIC_API_KEY:?Set ANTHROPIC_API_KEY in your environment}"
# SYSTEM_DOC should be a large, stable prefix (instructions + docs).
# Opus 5 caches from 512 tokens; pad or load a real file until usage shows a write.
SYSTEM_DOC="$(cat <<'SYS'
You are a repo assistant. Follow these rules on every turn:
1) Prefer existing helpers over new files.
2) Cite paths when you propose edits.
3) Keep answers short unless asked for a patch.
Project context (stable; do not invent APIs):
[PASTE a long static README, style guide, or OpenAPI excerpt here so the
cached prefix exceeds 512 tokens for claude-opus-5. For claude-sonnet-5,
target at least 1,024 tokens.]
SYS
)"
curl https://api.anthropic.com/v1/messages \
-H "content-type: application/json" \
-H "x-api-key: ${ANTHROPIC_API_KEY}" \
-H "anthropic-version: 2023-06-01" \
-d "$(jq -n \
--arg system "$SYSTEM_DOC" \
--arg user "List the three rules, then name one file you would open first." \
'{
model: "claude-opus-5",
max_tokens: 512,
tools: [
{
name: "get_file_outline",
description: "Return headings for a repo file path",
input_schema: {
type: "object",
properties: { path: { type: "string" } },
required: ["path"]
}
},
{
name: "search_repo",
description: "Keyword search across the indexed tree",
input_schema: {
type: "object",
properties: { query: { type: "string" } },
required: ["query"]
},
cache_control: { type: "ephemeral" }
}
],
system: [
{
type: "text",
text: $system,
cache_control: { type: "ephemeral" }
}
],
messages: [
{ role: "user", content: $user }
]
}')"
Automatic caching (top-level "cache_control": {"type": "ephemeral"}) is fine for growing chat history. Prefer explicit breakpoints when tools and system stay fixed while the user turn changes every request. You get up to four breakpoints per request; automatic caching consumes one slot when combined with explicit markers.
Gotchas
- Minimum is silent. Under-threshold prefixes never error; verify with usage fields.
- Minima are not monotonic. Opus 5 is 512; Sonnet 5 is 1,024; Opus 4.6 and Haiku 4.5 are 4,096. Check the docs when you change model ids.
- Breakpoint order. Cache covers a contiguous prefix in order
tools→system→messages. Putcache_controlon the last identical block, not on varying content. - TTL. Type
ephemeraldefaults to 5 minutes and refreshes on each hit. Use"ttl": "1h"for a 1-hour write at higher write pricing. Longer-TTL breakpoints must appear before shorter-TTL ones in the same request. - Exact match. Any byte change at or before the breakpoint invalidates that entry (reordered tools, tweaked system text, different
tool_choice/ thinking / effort settings as documented).
Recommended AI prompt
Copy this paragraph into ChatGPT, Claude, Gemini, Grok, or whatever you use.
Help me instrument Claude prompt caching for a production agent on claude-opus-5 (512-token minimum). Tools and system stay fixed; the user turn changes every request. Design where to place up to four explicit cache_control: {"type": "ephemeral"} breakpoints in tools → system → messages order, how to verify cache_creation_input_tokens then cache_read_input_tokens across two calls within the 5-minute TTL, what to do if both stay 0, and how the plan changes if we switch to claude-sonnet-5 (1,024) or Haiku 4.5 (4,096). Keep it concrete and copy-paste ready.
Sources:
Anthropic Prompt caching · Claude Platform Prompt caching · Claude Opus 5 overview