Claude Corner: Thinking Overspend → Set Adaptive Effort

  • Drop legacy budget_tokens. On Claude Opus 5, adaptive thinking is the default and manual thinking.type: "enabled" returns 400.
  • Steer depth with output_config.effort (low / medium / high / xhigh / max). Keep that field off the thinking object.
  • Bound cost with a large enough max_tokens (hard cap on thinking plus text) and read usage.output_tokens_details.thinking_tokens.

Your agent loop still burns tokens on trivial confirms because every turn inherits the same deep-reasoning posture. Fixed thinking budgets either overspend or 400 on current models. The Messages API expects adaptive thinking plus an explicit effort level instead.

Why it matters

Claude Opus 5 runs adaptive thinking by default: the model decides whether and how much to think per request. Effort is soft guidance for that decision. The API default is high, which is fine for hard coding turns and expensive for chatty tool confirmations. Legacy thinking: {type: "enabled", budget_tokens: N} is rejected on Opus 5 and later. Put the level on output_config.effort, leave room in max_tokens for reasoning plus the answer, and keep the same top-level effort across a cached conversation. Changing effort between requests invalidates prompt-cache breakpoints. The one move: set effort per workload class, raise max_tokens when you see stop_reason: "max_tokens", and watch thinking_tokens in usage.

How to set adaptive effort

Use host https://api.anthropic.com/v1/messages, model claude-opus-5, and secrets from the environment. No beta header is required for top-level effort. This tip uses output_config.effort only. Tip 3's output_config.format JSON schema is a different field on the same object.

  1. 1. Remove any budget_tokens / thinking.type: "enabled" payload for Opus 5+. Optional: set thinking: {"type": "adaptive"} (equivalent to omitting it on Opus 5).
  2. 2. Set output_config.effort to match the turn class: low or medium for routine steps, high (default) for complex work, xhigh or max for long agentic coding.
  3. 3. Size max_tokens for thinking plus visible text. At xhigh / max, start large (docs often suggest on the order of 64k) and tune from evals.
  4. 4. After HTTP 200, inspect stop_reason and usage.output_tokens_details.thinking_tokens. If you hit max_tokens, raise the cap or lower effort.
#!/usr/bin/env bash
set -euo pipefail
: "${ANTHROPIC_API_KEY:?Set ANTHROPIC_API_KEY in your environment}"

# Adaptive thinking (default on claude-opus-5) + soft effort steer.
# Do NOT put effort inside thinking. Do NOT send budget_tokens on Opus 5+.
curl -sS https://api.anthropic.com/v1/messages \
  -H "content-type: application/json" \
  -H "x-api-key: ${ANTHROPIC_API_KEY}" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
    "model": "claude-opus-5",
    "max_tokens": 16000,
    "thinking": {"type": "adaptive"},
    "output_config": {"effort": "medium"},
    "messages": [
      {
        "role": "user",
        "content": "List three trade-offs of caching HTTP responses at the edge. Keep it short."
      }
    ]
  }'
# Check stop_reason. If max_tokens, raise max_tokens or drop effort.
# Bill from output_tokens; thinking_tokens in output_tokens_details is the reasoning slice.
# Hold top-level effort steady inside a prompt-cached session.

Gotchas

  • Do not pass "adaptive" as an effort value. Adaptive is a thinking mode. Effort levels are low, medium, high, xhigh, and max.
  • On Claude Opus 5, thinking: {"type": "disabled"} with effort xhigh or max returns 400. Prefer lower effort with thinking left on.
  • Changing top-level output_config.effort between turns busts prompt-cache prefixes. For mid-conversation changes on supported models, use the per-message beta path instead of flipping the top-level field every turn.
  • Effort steers thinking volume more than visible reply length on Opus 5. Prompt for brevity when you need short prose.

Copy this paragraph into ChatGPT, Claude, Gemini, Grok, or whatever you use.

Sources: Effort · Thinking steering and cost · Extended thinking · Thinking overview

1 thought on “Claude Corner: Thinking Overspend → Set Adaptive Effort”

Leave a Comment