Claude Corner: Claude Code Quota Lockout → Cap Effort and Burn Fewer Tokens

  • Cap team burn with maxEffortLevel (Claude Code v2.1.267+) so routine sessions cannot climb into xhigh / max by accident.
  • Steer per session with /effort, --effort, or CLAUDE_CODE_EFFORT_LEVEL, and save defaults with effortLevel / modelSettings.
  • Watch the shared seat window with /usage, pick a cheaper model for routine work, plan first, and keep tool dumps short.

Your Team or Enterprise seat shares a rolling five-hour window and a weekly window across Claude Code, Claude chat, and Cowork. One long high-effort Claude Code loop that reads half the repo and fans out tools can empty that weekly allowance before lunch. The lockout is blunt: session or weekly limit hit, wait for the reset. Switching models does not restore a shared seat window.

Why it matters

Claude Corner: Thinking Overspend → Set Adaptive Effort covered adaptive effort on the Messages API. Teams that live in Claude Code feel the same dial as quota pain, not API request fields. Effort here is not only "think longer." Anthropic says it also shapes how many files Claude reads, how many tools it runs, and how far a multi-step task goes before checking in, so higher effort can mean many more output tokens for the same prompt. Seat allowances are shared across models, and Opus at xhigh on busy repos burns everyone faster. The one move: set a ceiling with maxEffortLevel, default routine work to medium or low, reserve high / xhigh for hard bugs, and treat /usage like a fuel gauge.

How to cap Claude Code effort (and other token leaks)

Use Claude Code settings under ~/.claude/settings.json (user) or .claude/settings.json (project). Documented keys: effortLevel, maxEffortLevel, modelSettings, model, bashOutputMaxChars, taskOutputMaxChars. Launch override: --effort. Env override: CLAUDE_CODE_EFFORT_LEVEL (beats both). Interactive: /effort and the effort slider in /model. Desktop secondary: the usage ring next to the model picker tracks the same plan story.

  1. 1. Run /usage before a heavy agent loop. Session and weekly seat windows are shared across models. An Opus-only limit is different and can be worked around with /model.
  2. 2. Cap the ceiling (v2.1.267+). Set "maxEffortLevel": "medium" or "high" in project or managed settings. Lowest cap across scopes wins. Devs can go lower; they cannot raise past the cap.
  3. 3. Default models without a saved per-model level with "effortLevel": "medium". Pin spendy models under modelSettings when needed.
  4. 4. For one cheap session: claude --effort low or export CLAUDE_CODE_EFFORT_LEVEL=medium. Inside a session, /effort medium (Claude Code reports whether it saved to modelSettings or applied session-only).
  5. 5. Add documented hygiene: /clear between unrelated tasks, plan before big edits (claude --permission-mode plan or Shift+Tab), keep CLAUDE.md lean, disable idle MCP servers, trim Bash/task dumps with bashOutputMaxChars / taskOutputMaxChars, and leave ultracode off unless you truly want xhigh plus workflow planning.
mkdir -p .claude
cat > .claude/settings.json <<'EOF'
{
  "model": "sonnet",
  "effortLevel": "medium",
  "maxEffortLevel": "high",
  "modelSettings": {
    "claude-opus-5": {
      "effortLevel": "medium",
      "maxEffortLevel": "high"
    }
  },
  "bashOutputMaxChars": 4000,
  "taskOutputMaxChars": 4000,
  "autoCompactEnabled": true,
  "ultracode": false
}
EOF

# Flag beats effortLevel; CLAUDE_CODE_EFFORT_LEVEL beats both
# export CLAUDE_CODE_EFFORT_LEVEL=low
claude --effort low --permission-mode plan

# When you need more depth later:
# /effort high
# /model    (effort slider for active model)
# /clear    (drop stale context between unrelated tasks)

Gotchas

  • Effort steers token use; it is not a hard budget. Soft controls are the Claude Code levers. The API-side hard truncate remains max_tokens.
  • maxEffortLevel needs v2.1.267+. A value of "max" means that source places no ceiling.
  • ultracode / --effort ultracode runs at xhigh. A maxEffortLevel below xhigh turns ultracode off for that model.
  • Changing effort mid-session can reprice the next turn because prompt cache keys by effort as well as model. Prefer setting effort at session start on tight seats.
  • Session and weekly seat limits are shared across models. Only model-family limits clear when you switch families with /model.

Copy this paragraph into ChatGPT, Claude, Gemini, Grok, or whatever you use.

Sources: Claude Code effort + model selection · Manage costs · Settings reference · Model config · Errors

1 thought on “Claude Corner: Claude Code Quota Lockout → Cap Effort and Burn Fewer Tokens”

Leave a Comment