secd3v · Claude Code · Developer Training

Developer Cost
Optimisation Training

Twelve developer behaviours control what Claude Code costs you, your team and your organisation. Each factor states its cost impact, explains the mechanism, gives worked examples from DevSecOps, agentic and spec-driven work, and ends with a knowledge check. Every figure is priced on Amazon Bedrock Australian endpoints with the 10% sovereignty premium applied.

REGION=AU · PROFILES=E8 ML2 or PROTECTED · v4.0
How to use this guide. Start with Working Securely. Permission mode behaviour frames every session you run, and the optimisation factors assume you know what the classifier does and what it leaves to you. Then work the factors in order. Factor 01 carries the highest return of the twelve and takes less than a day to internalise. Factor 04 changes how you plan a whole sprint. The remainder protect the return those two produce rather than adding to it. Reference is a lookup, not a read: it carries the model list, how each one caches, why token counts differ between generations, and the planning allowance for work the cost figures exclude.

Before the factors

Foundations
Working Securely
Permission modes, untrusted content, transcripts
Foundations
Prerequisites
The CLI on your PATH gates six behaviours
Foundations
The Three Work Patterns
DevSecOps, agentic, spec-driven

The twelve optimisation factors

Factor 01
Prompt Specificity
Highest · All patterns
Factor 02
Plan Mode
Highest · Agentic
Factor 03
Verification Targets
High · All patterns
Factor 04
Spec-Driven Development
Highest · Agentic · 28% at heavy
Factor 05
Session Hygiene
High · DevSecOps
Factor 06
Model Selection & Effort
High · DevSecOps
Factor 07
Extended Thinking
High cost when unmanaged
Factor 08
Subagents
High · Agentic
Factor 09
Preprocessing Hooks
High · DevSecOps
Factor 10
CLAUDE.md, Skills & Context
Medium · All patterns
Factor 11
CLI Tools & Plugins
Medium · All patterns
Factor 12
Token Telemetry
Foundation for everything above

What a developer costs

Standard developer policy with no Opus access, per month in USD, 22 working days, Australian endpoint rates. Use these to recognise when your own telemetry has drifted.

$65.62DevSecOps heavy, no MCP
$151.49Agentic heavy, no MCP
$108.76Spec-driven heavy, no MCP Est

Prompt caching is the largest single lever, cutting cost by 38 to 63% against an uncached baseline. The 5-minute default captures that saving with no configuration and no premium. Spec-driven development is the second, cutting heavy agentic spend by 28% for the cost of one authoring session per sprint.

Add 10% for work these figures exclude. Every figure models software development: reading and writing code, reviewing changes, generating tests, auditing for defects, building against a specification. Long-form documentation, release notes, merge request summarisation, diagrams, commit messages, incident write-ups, runbooks and onboarding material are not costed above. One such session per day adds roughly 8% at a heavy DevSecOps tier and 3% at a heavy agentic one. Reference carries the allowance and when to use 5%, 10% or 15%.
Start here
Working Securely
→
Wiki / Foundations / Working Securely
Foundations · Read first

Working Securely

The secd3v Claude Code Service implements ISM PROTECTED controls and can support unclassified through to PROTECTED workloads. Customers can deploy the Claude Code CLI with CCS using a range of different security profiles, for Essential Eight Maturity Level 2 (Guide 1) or ISM PROTECTED (Guide 2). Most of the posture is managed configuration you cannot change. This page covers what depends on you.

TL;DR
Use auto mode. It is the recommended default on this deployment and is pinned for you through managed settings. Auto mode does not remove the safety gate, it moves it: your ask rules still fire first, then a classifier reviews everything below them. The actions your organisation defined as risky continue to stop and ask by name. Reviewing what the agent proposes before it lands is the one duty that never automates. Start sessions from the repository root, and assume your open editor file reaches the model on every turn.
How to read this page. Everything below is the Guide 1 baseline, which covers Essential Eight Maturity Level 2 work at OFFICIAL and OFFICIAL: Sensitive. Where Guide 2 tightens something for PROTECTED work, it is marked At PROTECTED. If you do not know which profile your team runs, ask your platform owner, because a few of the constraints below differ.

Permission modes: use auto mode

Claude Code sessions start in Manual mode, where Claude asks before file writes and most shell commands. Auto mode is opt-in on Amazon Bedrock and must be pinned through managed settings using defaultMode, so it is configured for you rather than something you switch on per session. Anthropic has stated an intention to make it the default on cloud platforms without committing to a date.

Auto mode is the recommended default on this deployment. The secd3v PROTECTED deployment guide permits it with organisationally defined risky actions held as ask rules, and Anthropic's published evaluation found the classifier more accurate than manual approval. The reason to prefer it is not convenience: constant prompting on trivia is what trains developers to click through everything, and approval fatigue is a security failure mode in its own right. Auto mode keeps the prompts rare so the ones that appear get real attention.

Auto mode does not remove the safety gate, it moves it

Instead of prompting on routine work, auto mode routes every tool call through a separate classifier that evaluates whether the action is irreversible, destructive, or aimed outside your environment. Data exfiltration sits in a hard-deny category the classifier is designed never to approve. Actions it judges safe proceed without interrupting you.

Permission rules fire before the classifier. That ordering is what matters. The ask rules implementing the ISM-2113 human approval gate over the organisationally defined risky set continue to hold in auto mode, because they are evaluated first and the classifier never sees those calls. Auto mode changes what happens to everything below that line, not the line itself.

Two gates, with different reach

GateCoversMechanism
Agent ask rulesCommands and file operations the agent runs on your endpoint, over the organisationally defined risky setManaged ask rules, evaluated before the classifier, binding in every permission mode
CCS GitLab MCP checkpointsHigh-risk platform operations, including merge approvals and mutations on the code platform, which the agent's own controls do not reachServer-side checkpoints, plus a default posture that suppresses mutating tools from the surface entirely

The gateway carries no human-in-the-loop mechanism of its own, so these two are the gates. They cover different ground rather than duplicating each other: an ask rule stops a command you run, and a checkpoint stops a platform operation the agent invokes through the MCP server.

One duty never automates. Review what the agent proposes before it lands. The classifier reduces noise so that duty gets real attention, and the curated allow lists keep it off trivia, but neither is a substitute for reading a change. --dangerously-skip-permissions is out of scope on this deployment and the service rejects it.
Manual mode costs more for the same work. Each approval round trip adds a turn, so a team that mandates Manual across the board should expect higher turn counts and higher monthly figures than the cost model shows, since every figure there assumes auto mode with ask rules.

What still asks you, and why it matters

Your ask and deny rules bind in every permission mode and are evaluated before the classifier, so the organisational risky action set below stops and asks by name regardless of what the classifier would have decided. This is the ISM-2113 human approval gate in practice.

ActionPrompts youProfile
git pushAlwaysGuide 1 (E8 ML2)
Package publicationAlwaysGuide 1 (E8 ML2)
Deletion outside the working treeAlwaysGuide 1 (E8 ML2)
Any command carrying a credentialAlwaysGuide 1 (E8 ML2)
Merge approvalsAlwaysGuide 2 (PROTECTED)
Pipeline retry and cancel operationsAlwaysGuide 2 (PROTECTED)
Any command moving data across the containment boundaryAlwaysGuide 2 (PROTECTED)

Every action on this list prompts you always, and the Profile column says which deployments carry it. The first four are the Guide 1 set that every deployment carries. The three shown in blue apply only where your team runs Guide 2, because at PROTECTED the containment boundary and the code platform both become things a human signs off rather than a classifier. Your team may have extended the set further, so treat the list as the floor rather than the whole of it.

Every prompt on this list deserves a read. Auto mode and the curated allow lists in version control exist so your everyday loop of test runners, linters and builds runs without interruption. That is precisely what makes the remaining prompts meaningful: a prompt reaches you because the action was deliberately excluded from classifier authority, not because it was cleared.

Untrusted content

You handle untrusted content every day through merge request diffs, issue text, log files, dependency source and pipeline output. Four practices apply.

  • 1Review suggested commands before you approve them. The prompt is the control, and skimming it defeats the control.
  • 2Do not pipe untrusted content directly to Claude. Preprocess it first. Factor 09 covers the mechanism, which cuts cost as well as risk.
  • 3Verify proposed changes to critical files. Migrations, infrastructure code and authentication logic warrant a read rather than a glance.
  • 4Report suspicious behaviour with /feedback.

The agent cannot fetch

The single highest-value deny rule on the baseline closes the most direct ingestion and exfiltration path for you. WebFetch, WebSearch, curl, wget and the PowerShell download commands are denied by managed rule. That removes the agent's fetch capability, not your endpoint's: your browser and your tooling are unaffected.

Mediated egress variant. Where your organisation runs an authenticated proxy with a domain allowlist, the deny rules become defence in depth and your team may permit WebFetch against allowlisted domains for developer experience. The Bash download denies stay either way, because shell downloads are harder to constrain by destination. Check with your platform owner before assuming a fetch will work.
At PROTECTED. Guide 2 removes the route rather than the capability. The agent runs inside a virtual machine or development container on your endpoint, with egress allowlisted to the service and your own infrastructure only, so no path to the public internet exists to deny. Packages come from internal mirrors rather than public registries, which means availability moves at the speed of your mirror curation. The deny rules stay in place anyway, turning a routing failure into a second refusal.

Detection sits behind the structural controls

Your organisation runs endpoint detection and response with AI agent runtime inspection, scanning the agent loop at your prompt, before each tool call, and on each tool response. In block mode a detected injection stops before the action runs, and you are notified both in the agent and by system notification. At PROTECTED, block mode is required rather than a maturity target, and the agent runs inside the contained environment where the inspection happens.

Every refusal has a request path rather than a workaround. A denied fetch, a server not on the approved list, or a blocked action with an EDR notification all mean the same thing: raise it, do not route around it. Working around a control is the failure the whole posture exists to prevent.

Boundaries, settings and credentials

Start from the repo root
Never from your home directory. The CLI writes only inside the directory it starts in and its subfolders, and outside auto mode asks before reading beyond it, so the boundary is only meaningful if the directory is. You confirm trust in a new codebase once.
Bash sandboxing
Shell commands run with filesystem and network isolation. A command that works in your own terminal may be refused inside a session, and that is the sandbox rather than a fault.
Managed settings win
Your posture lives in one administrator-owned file, managed-settings.json, which is read-only to you. It overrides anything you or a repository set, and managed deny rules cannot be overridden by a lower-tier allow rule. You cannot configure your way out of the posture, so where a rule blocks legitimate work, raise it through the request path rather than working around it.
Configuration surface is locked
A ConfigChange hook logs or blocks in-session settings changes, hooks are restricted to managed sources, and permission bypass is disallowed.
Deny-read fences
.env files, key files and credential directories. A Read deny on a sensitive path also prevents the editor selection and open-file notice for that file reaching the model, which serves content control and cost together.
MCP servers are deny by default
managed-mcp.json lists only the CCS GitLab MCP server, and the CLI loads nothing else: you cannot add servers, repository .mcp.json files are ignored, plugin servers are suppressed and --mcp-config is refused. An MCP server is remote code with tool authority over the agent, so it gets governed like software you install. A blocked add attempt shows an enterprise policy error, but a previously configured server that becomes blocked simply disappears without explanation. Propose a server through the request path rather than configuring around it.
GitLab runs under your identity
GitLab context arrives through the CCS GitLab MCP server under your own credentials, and its default posture suppresses mutating tools from the surface entirely. Any relaxation of that posture is a service decision, not a developer setting.
Transcripts
The service retains nothing, so your session transcripts under your user profile are the only place prompt content persists. They are your best investigation artefact and a data holding your organisation governs, with retention set through the CLI's cleanup setting. At PROTECTED they are a PROTECTED holding held inside the contained environment, so do not copy them out of it.

The editor context you are sending without noticing

When the CLI runs inside Visual Studio Code it connects automatically to the editor's in-process MCP server. While connected, it includes your current editor selection and the path of the active file as context on every prompt you send. The transcript shows a selected-lines notice when this happens. That context is variable, uncached and outside every budget figure in this guide. Close what you are not working on, and use a Read deny rule for any path that should never reach the model.

This one server is the exception to the deny-by-default MCP position above. The extension's in-process server still loads in sessions the extension starts, regardless of managed-mcp.json, because it is product behaviour rather than a configurable setting. At PROTECTED the editor runs connected into the contained environment through remote-development tooling, which makes the boundary invisible in the editor but does not change this behaviour.

Security tooling available to you

/security-review

Runs an on-demand security pass over the changes on your current branch. Worth running before any merge request touching authentication, input handling or data access.

Security guidance plugin

Has Claude review and fix vulnerabilities in its own code changes during the session, rather than waiting for a review gate to find them.

Overall, the posture holds without your attention in most respects, and the two places it does not are the prompts you approve and the content you feed in.

Knowledge Check · Working Securely
Scenario: You are running a long agentic session in auto mode. Claude proposes a bash command that deletes a build directory outside your working tree, and you receive a prompt. You are eight turns into a flow and want to keep moving. What is the right read of that prompt?
A. Approve it. The classifier already reviewed it, so the prompt is a formality
B. The prompt exists because deletion outside the working tree is on the organisational risky action set. It reached you precisely because it is not a formality. Read what is being deleted before approving
C. Switch to Manual mode so you get more prompts
D. Deny it and restart the session
Show answer
B is correct. Ask rules are evaluated before the classifier, so this action never reached it. A inverts the logic: the prompt appears because the action was deliberately excluded from classifier authority, not because it was cleared. C moves you off the recommended default and buys noise rather than control, and noise is what trains developers to click through. D discards eight turns of context for a question a five-second read answers.
←
Previous
Overview
Next
Prerequisites
→
Wiki / Foundations / Prerequisites
Foundations

Prerequisites

Running claude in the integrated terminal requires the standalone CLI on your shell PATH. Installing the graphical extension does not provide it, because the extension bundles a private copy of the CLI for its own chat panel.

The CLI installs to the managed, administrator-owned path the deployment guides mandate, and that path is added to the system PATH through the managed environment configuration rather than your user profile. The application control position does not change. Verify with claude --version before your first session.

Without the PATH entry, six behaviours this guide teaches are unavailable, and each is a cost-reducing behaviour elsewhere in the training.

CommandPurposeTaught in
claudeLaunching the CLI at all, which is the supported configurationEverywhere
claude --permission-mode planStarting directly in plan modeFactor 02
claude -pNon-interactive runs, including pre-commit hook automation and fan-out loops over a file listFactor 09
claude --resume, --continueReturning to a named session instead of rebuilding contextFactor 05
claude mcp listConfirming which MCP servers a session will load, the managed MCP validation checkFactor 10
claude --versionVersion verificationFactor 12
One limit applies to the automation cases. Machine-to-machine access is recorded as an open gap. A CI job reaches the CCS GitLab MCP server tool surface but not the gateway's inference path, so a pipeline that needs inference is not currently served. Non-interactive inference is available under your own session on your own workstation. Plan pre-commit hooks and local fan-out loops with claude -p, and do not plan a pipeline stage that calls the model.

/clear and /compact are in-session commands and do not depend on PATH, although starting the CLI does. If claude --version fails, raise it with your platform owner rather than installing the CLI yourself, since a user-installed copy will not satisfy application control.

←
Previous
Working Securely
Next
The Three Work Patterns
→
Wiki / Foundations / Work Patterns
Foundations

The Three Work Patterns

Two foundational patterns cover secd3v Claude Code usage, with a third mode that modifies how one of them runs. Identify which one you are in before you apply any factor, because the same behaviour carries different value in each.

PatternWhen it appliesSession shapeDominant cost driver
DevSecOpsExisting production system with live users. Code review, bug fixes, security checks, merge request review under active security constraints8 to 11 short sessions per day, 5 to 45 min each. You approve every stepCache reads across many short sessions, plus regular input from broad prompting
AgenticBuilding something new. New service, new module, greenfield implementation where wrong output gets discarded cheaply1 to 2 long sessions per day, up to 4 hours. Claude works autonomouslyRegular input at 25,000 tok per turn at heavy usage, 4.1 times the DevSecOps volume
Spec-DrivenAgentic variant. Author a specification before executing, then execute against it rather than discovering scope through file explorationPhase 1 spec write 30 min, Phase 2 execution up to 4 hrs, Phase 3 conformance review 30 minRegular input cut to approximately 13,000 tok per turn, a 48% reduction

Why DevSecOps is the brownfield pattern

Brownfield work means an existing production system with live users, an established architecture, years of accumulated code, technical debt and security constraints under active enforcement. You cannot start fresh, cannot discard a wrong implementation without consequence, and cannot hold broad autonomous permissions without risk. In government, defence and high compliance contexts this describes the overwhelming majority of daily work.

DevSecOps is the direct response to those constraints rather than a methodology that happens to sit alongside them. Where a codebase has a security posture that cannot be accidentally degraded and a production environment where mistakes have immediate consequences, the human-gated, review-at-every-step workflow becomes mandatory rather than optional discipline.

That shapes how you use the CLI. You say "look at auth.py lines 42 to 89 and identify any SQL injection risk" rather than "explore this codebase and make the changes you think are needed". Sessions run short, targeted and bounded, because the work demands precision over autonomy.

Why agentic is primarily the greenfield pattern

Greenfield work builds a new service, module or application, with no live users to disrupt, no established architecture to break and no accumulated constraints to misunderstand. A wrong-direction implementation gets discarded cheaply, so the cost of an incorrect autonomous attempt stays low. Claude explores, plans and implements, reading many files and building its own context map over long sessions.

Agentic development also suits specific brownfield scenarios, including large-scale migration sprints, comprehensive test generation and automated documentation. Those cases require three conditions: plan mode before any execution, a git worktree for isolation, and every checkpoint treated as a safety gate.

Agentic development does not suit routine brownfield maintenance, security operations, or any work where an incorrect autonomous change reaches production. The session risk profile determines the fit, not the label on the codebase.

Session shapes

Use these to recognise when a session has drifted out of its intended type. Factor 12 turns them into telemetry thresholds.

PatternTypeShape and typical workTurnsRegular input / turnOutput / turn
DevSecOpsMicro5 turns, 5 to 10 min. Doc update, syntax check, single-function review, pipeline triage, quick explanation52,500 tok400 tok
DevSecOpsStandard9 turns, 15 to 25 min. Code review, targeted bug fix, test generation, single-file security check, MR feedback95,000 tok700 tok
DevSecOpsExtended15 turns, 30 to 45 min. Multi-file security audit, SAST triage, compliance check, refactoring plan and execute159,000 tok1,000 tok
AgenticLight5 sessions of 10 turns. Small feature additions, single-module builds, focused bug fixes50 / day4,000 tok600 tok
AgenticMedium2 sessions of 25 turns. Feature implementation, module construction, MR creation50 / day13,000 tok800 tok
AgenticHeavy1 session of 65 turns plus 1 of 25. New service construction, large autonomous implementation from spec90 / day25,000 tok1,100 tok

Spec-driven sprint days

A spec-driven sprint runs three day shapes rather than one. Recognising which day you are in tells you which model and effort level belong.

DaySession mixTurnsWhat happens
Spec day1 spec-write, 1 execution start~25Phase 1 plus Phase 2 kickoff. Spec authored, execution begins
Exec day1 to 2 execution sessions65 to 90Phase 2 sustained. Building against the spec, /compact as needed
Review day1 conformance, 1 correction~20Phase 3 review. Deviation notes feed a Phase 2 correction turn
These shapes assume auto mode with ask rules, which is the recommended default. Manual mode raises turn counts for the same work, because each approval round trip adds a turn, so a team that mandates Manual should expect higher turn counts and higher monthly cost than the cost model shows. See Working Securely for why auto mode does not weaken the approval gate.

Which factors pay in which pattern

DevSecOps developers
  • Factor 01 · Prompt specificity
  • Factor 05 · Session hygiene
  • Factor 06 · Model selection and effort
  • Factor 09 · Preprocessing hooks
Agentic developers
  • Factor 02 · Plan mode
  • Factor 03 · Verification targets
  • Factor 04 · Spec-driven development
  • Factor 08 · Subagents

Overall, the pattern determines the dominant cost driver: DevSecOps spends on cache reads across many short sessions, while agentic spends on regular input inside a small number of long ones.

←
Previous
Prerequisites
Next
Factor 01 · Prompt Specificity
→
Wiki / Optimisation / Prompt Specificity
Factor 01 of 12

Prompt Specificity & Context Front-Loading

The largest per-session cost lever you control. How you phrase a request determines how much context Claude reads before it can act, and every token read compounds across the rest of the session.

TL;DR
Name the file, the lines and the concern. A prompt carrying an explicit path, line range and specific concern costs three to five times less than a vague equivalent, because Claude reads only what you point to rather than everything that might be relevant.
💰
Cost impact: highest, all patterns
Claude Code's context is cumulative, and every API call processes the full conversation history to date. A vague prompt makes Claude read broadly to find relevant files, and those reads stay in context for every subsequent message.
3 to 5 times regular input reduction
~40,000Tokens from broad file exploration triggered by one vague prompt on a medium module
~8,000Tokens for the same task with an explicit file reference and line range
5×Reduction in regular input from specific against vague prompting on this task type

DevSecOps examples

Security review of an authentication function

✗ Vague
Check my auth code for security issues
Claude scans auth.py, user.py, middleware.py, session.py, config.py, database.py and models.py, then their imports. Context grows before the work begins.
✓ Specific
Review @src/auth/login.py lines 42-89 for SQL injection risk. The database layer uses psycopg2 and our connection is in @src/db/connection.py. Flag any raw string interpolation into queries.
Two files. Claude knows what to look for and where, and starts work immediately.

Fixing a known bug

✗ Vague
The user profile API is returning 500 errors. Fix it.
No file, no error message, no environment. Claude reads the module and its dependencies, potentially ten or more files, before forming a hypothesis.
✓ Specific
@src/api/user_profile.py line 134 throws: KeyError: 'preferred_name' The field was added to the User model last sprint but the serializer at line 134 still expects the old schema. Fix the serializer to handle the field being absent, defaulting to None.
One file, exact line, exact error, exact expected behaviour. A targeted edit inside a micro session.

Merge request review with a security focus

✗ Vague
Can you review this merge request for me?
No MR number, no focus area, no standard. Claude asks clarifying questions, which wastes turns, or reads the whole diff plus related files speculatively.
✓ Specific
Review MR !847, which adds JWT token refresh to session middleware. Focus on: 1. Token expiry edge cases, expired-during-request 2. Whether refresh tokens are invalidated on logout 3. Any timing window allowing token reuse Session middleware is at @src/middleware/session.py
Three specific concerns, one target file. Focused analysis costs far less than a general sweep.

Agentic example: a new API endpoint

✗ Vague
Add an endpoint to the API for managing notifications.
Claude reads the API module, models, routing config and tests to work out the patterns before writing anything. Those reads stay in context for 60 or more subsequent turns.
✓ Specific
Add GET /api/v2/notifications following the pattern in @src/api/v2/messages.py. - Paginated, use PaginationMixin from @src/api/mixins.py - Filter by ?status=read|unread - Auth via the existing @require_auth decorator - Return: id, title, body, created_at, is_read The Notification model already exists, no new models needed.
Two reference files. Claude implements immediately, saving ten or more exploration turns at the start of a long session.

The 5W1H checklist

Answer these six before sending. If you can answer them, write the answers into the prompt.

QuestionWhat it coversExample in prompt
WhatExactly what to change, review or understand"review the token validation logic" not "check auth"
WhereSpecific file path and line numbers"@src/auth/tokens.py lines 88-134"
WhyError message, security concern, test failure"returning 401 for valid tokens after migration"
Look forSpecific patterns or vulnerability class"flag non-constant-time string comparisons"
Don't touchFiles or logic to leave unchanged"do not modify the database schema or migrations"
HowPattern or library to follow"use the same approach as @src/auth/refresh.py line 44"
One task per prompt. Bundling three tasks, such as "fix the bug, add tests and update the README", forces Claude to hold three dependency trees in context at once. Sending them separately costs less in total even though it is more messages.

Specificity works differently in each pattern

PatternWhat specificity does
DevSecOpsKeeps sessions inside their session-type scope and stops a micro session drifting into extended territory
AgenticVague prompts trigger broad file scanning, and each read compounds into every subsequent turn of a long session
Spec-DrivenThe spec is the specificity mechanism, but the execution prompt must still reference specific spec sections rather than leaving Claude to interpret the whole document

Quick reference by task type

TaskInstead ofSay
Security check"check for security issues""review @api/views.py lines 55-90 for IDOR, ensure user_id is validated against request.user"
Test generation"write tests for the payment module""write pytest unit tests for charge() in @payments/processor.py using fixtures from @tests/conftest.py, cover success, declined card, network timeout"
Documentation"document the API""add Google-style docstrings to the 4 public methods in @src/api/client.py, type annotations already present, skip them"
Pipeline triage"CI is failing, fix it""job 'test-unit' in pipeline #4821 fails with ImportError: cannot import 'TokenCache' from 'src.cache', class renamed last commit. Fix imports in test files only"
Refactoring"refactor the database layer""extract retry logic from @db/connection.py lines 120-156 into a RetryPolicy class in @db/retry.py, keep the existing interface, callers must not change"

When a vague prompt is the right call

A vague prompt is useful when you are exploring and can absorb the course correction. What would you improve in this file? surfaces things you would not have thought to ask about. Use it deliberately, in a session you intend to clear, rather than as a default.
Knowledge Check · Factor 01
Scenario: You need to review a payment function for timing attack vulnerabilities. You know it sits in src/payments/processor.py but not the exact lines. Which prompt costs least?
A. "Review the payment code for security issues"
B. "Check src/payments/ for timing vulnerabilities"
C. "Review the compare_secret() and validate_token() functions in @src/payments/processor.py for timing attack vulnerabilities, look for non-constant-time string comparisons"
D. "Audit the entire payments module for OWASP Top 10 issues"
Show answer
C is correct. Naming the specific functions and the exact vulnerability class limits scope to one file and one issue type even without line numbers. A and D scan entire modules. B scans a directory. C is precise enough to produce a targeted, affordable session.
←
Previous
The Three Work Patterns
Next
Factor 02 · Plan Mode
→
Wiki / Optimisation / Plan Mode
Factor 02 of 12

Plan Mode Before Execution

Plan mode separates what Claude should work out from what it should build. Correcting direction at the plan stage costs roughly 500 tokens. Correcting after twenty turns of wrong implementation costs tens of thousands.

TL;DR
For any task touching multiple files, or where you are uncertain of the approach, press Shift+Tab before typing. There is no /plan command. Review the plan, edit it with Ctrl+G, then switch back to execute. Skip plan mode for small clearly scoped tasks where you could describe the diff in one sentence.
💰
Cost impact: highest agentic. Mandatory for multi-file brownfield changes
In plan mode Claude reads files and runs shell commands to explore, and makes no edits to your source. The exploration costs tokens, and far fewer than an incorrect implementation discovered at turn twenty.
~500 tokens to correct a plan, tens of thousands to correct an implementation

How to enter and leave it

  • 1Enter. Press Shift+Tab until the status bar shows plan mode on, or start with claude --permission-mode plan. The start flag depends on the CLI being on your PATH.
  • 2Explore. Claude reads relevant files and asks clarifying questions without touching your code.
  • 3Review and edit. Press Ctrl+G to open the plan in your text editor. Correct the approach, add constraints, or reject it, all at near-zero cost.
  • 4Execute. Approve the plan or press Shift+Tab to switch back. Claude implements against the plan it agreed.

When to use it

✓ Use plan mode when
  • The task touches multiple files or modules
  • You are unfamiliar with the code being changed
  • You are uncertain what the right approach is
  • Any brownfield change to a production system
  • New feature implementation in an existing service
  • Refactoring with potential side effects
  • Security-sensitive changes
  • Spec-driven Phase 2, before each significant component
○ Skip plan mode when
  • You could describe the complete diff in one sentence
  • Fixing a typo, renaming a variable, adding a log line
  • Adding a docstring to a clearly understood function
  • Updating a dependency version
  • A formatting-only change
  • Any task where the scope is fully clear upfront
Plan mode adds overhead on trivial tasks. Anthropic's test is the one to use: if you could describe the diff in one sentence, skip the plan.

DevSecOps example: multi-file bug fix

✗ Without plan mode
Fix the password reset flow. Users aren't receiving the reset email and the link expires too quickly.
Claude dives into implementation. It may fix email delivery in the wrong layer, change token expiry in a way that conflicts with session middleware, and miss a third issue in validation. You discover this at turn 15, and rewinding costs the whole investment.
✓ With plan mode
[Shift+Tab to plan mode] Fix the password reset flow. Users aren't receiving the reset email and the link expires too quickly.
Claude maps the flow across auth/reset.py, notifications/email.py and middleware/session.py, identifies all three issues, and proposes a plan. You accept the email fix and correct the session expiry approach before a line is written.

Agentic example: new service construction

✗ Without plan mode
Build a notification service that sends email, SMS and push notifications. It should be async and support retry logic.
At turn 30 you discover the queue is built on Redis when your infrastructure uses SQS, and the retry logic does not match your circuit-breaker pattern. Starting over costs the entire 30-turn investment.
✓ With plan mode
[Shift+Tab to plan mode] Build a notification service that sends email, SMS and push notifications. Async with retry logic. We use SQS for queuing (@infrastructure/queues.py) and our circuit-breaker pattern is in @src/resilience/circuit_breaker.py
Claude reads your infrastructure files and proposes an architecture that fits. You catch the missing dead-letter queue requirement and add it to the plan. Implementation proceeds correctly from turn 1.
Plan mode in DevSecOps is a compliance practice, not just a cost practice. On brownfield production systems, executing before planning is how you modify the wrong component or introduce a change that passes tests and carries a subtle security regression.
Plan mode and spec-driven development operate at different granularities. Plan mode prevents wrong-direction turns inside a single session. A spec prevents wrong-direction sessions from starting at all. Both are essential. See Factor 04.
Knowledge Check · Factor 02
Scenario: You need to add input validation to a single function in @src/api/users.py. The logic checks that the email field matches a regex. Use plan mode?
A. Yes, always use plan mode for any API change
B. No. Single function, clear scope, one-sentence diff. Ask Claude directly
C. Yes, any change to a brownfield file needs a plan
D. Only if the function exceeds 50 lines
Show answer
B is correct. One function, one field, one regex check. This is exactly the case Anthropic describes as skip the plan. Reserve plan mode for genuine scope uncertainty or multi-file impact.
←
Previous
Factor 01 · Prompt Specificity
Next
Factor 03 · Verification Targets
→
Wiki / Optimisation / Verification Targets
Factor 03 of 12

Verification Targets

Claude stops when the work looks done. Without a check it can run, "looks done" is the only signal available, and you become the verification loop. That is the most expensive loop in the workflow, because every mistake waits for you to notice it.

TL;DR
Give Claude something that returns pass or fail: a test suite, a build exit code, a linter, or a script that diffs output against a fixture. The loop then closes without your attention. Ask for evidence rather than assertion, meaning the test output and the command run.
💰
Cost impact: high, all patterns
A session you have to verify by hand is a session that cost you attention as well as tokens, and attention is the scarcer resource.
Closes the loop without developer attention

What counts as a check

Anything returning a signal Claude can read in the conversation. A test suite, a build exit code, a linter, or a script that diffs output against a fixture.

Instead ofWrite
"implement a function that validates email addresses""write a validateEmail function. Test cases: user@example.com is true, invalid is false, user@.com is false. Run the tests after implementing"
"the build is failing""the build fails with this error: [paste error]. Fix it and verify the build succeeds. Address the root cause, do not suppress the error"
"add retry logic to the client""add retry logic to @src/http/client.py. Write a test asserting three attempts on a 503 then a raise. Run the suite and show me the output"
"make the dashboard look better""[paste screenshot] implement this design. Screenshot the result, compare it to the original, list the differences and fix them"

Ask for evidence, not assertion

✗ Assertion
I've fixed the issue. The logic now handles the null case correctly.
Plausible, unverified, and you now have to run the suite yourself to find out. You are the verification loop.
✓ Evidence
Ran: pytest tests/api/test_profile.py -q Result: 14 passed, 0 failed The KeyError came from the serializer at line 134. Added a .get() with a None default and a regression test at test_profile.py:88.
Reviewing evidence costs less than re-running the verification yourself, and it works for sessions you were not watching.

Three gate strengths, and what each costs

GateMechanismCost
In-promptAsk Claude to run the check and iterate in the same messageNothing beyond the turns the iteration takes. Use this by default
/goal conditionA separate evaluator re-checks the goal after every turn until it resolvesThe evaluator runs every turn. Idle check-ins start a new turn carrying full context while background work runs. CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 is set unless goal-driven sessions are approved for your team
Stop hookA script blocks the turn from ending until the check passesUp to 8 consecutive blocks before Claude Code overrides the hook. Budget up to 8 extra turns per gated task, and keep the check cheap and fast
Each step trades setup for attention. Start with the in-prompt version: it works on any task today and costs nothing to configure. Move to /goal or a Stop hook only where a run needs to finish correctly without you watching it.

The second opinion

A verification subagent reviewing the diff in a fresh context is the cheapest quality gate available, because the agent doing the work is not the one grading it. Factor 08 covers the pattern and the prompt.

Knowledge Check · Factor 03
Scenario: You ask Claude to fix a failing integration test and walk away. You come back to "I've fixed the issue, the logic now handles the null case correctly." What went wrong in how you set the task up?
A. Nothing, that is a normal completion message
B. You gave no verification target. Claude stopped when the work looked done, and you now have to run the suite yourself to find out whether it did
C. You should have used Opus 4.8
D. You should have used plan mode
Show answer
B is correct. The fix may well be right, but you have an assertion rather than evidence, and you are the verification loop. The prompt should have ended with "run the integration suite and show me the output". Neither a larger model nor plan mode addresses the missing check.
←
Previous
Factor 02 · Plan Mode
Next
Factor 04 · Spec-Driven Development
→
Wiki / Optimisation / Spec-Driven Development
Factor 04 of 12

Spec-Driven Development

The largest single return available in the cost model. Author a complete specification before execution begins, and cut heavy agentic spend by 28% for the price of one authoring session per sprint.

TL;DR
Before any agentic execution, write a specification: interface contracts, data shapes, acceptance criteria, file layout, security constraints, test scope. Save it to SPEC.md, end the authoring session, then execute against it in a fresh session. Claude then implements rather than discovering scope through file exploration.
💰
Cost impact: highest agentic. Approximately 28% monthly saving at heavy Est
Standard agentic heavy sessions average 25,000 tokens regular input per turn, driven by file exploration and growing conversation history. Spec-driven execution averages approximately 13,000, a 48% reduction. Wrong-direction turns fall from four to eight per session to zero or one. At heavy usage without MCP, monthly spend drops from $151.49 to approximately $108.76, a saving of $42.73.
~48% regular input reduction · Phase 3 Haiku-eligible · Opus justified at Phase 1 for all tiers
The saving concentrates at heavy usage. At light usage it falls to approximately 7%, because a light agentic session does little of the file exploration a spec removes. All spec-driven figures are estimated and will be restated at the 30-day telemetry review. They model Phase 2 regular input at 13,000 tokens per turn against 25,000 for standard agentic, include the Phase 1 Opus overhead at one spec session per five execution sessions, and assume Phase 3 runs on Haiku. Actual results depend on spec quality and sprint cadence.

Why it matters beyond cost

Spec-driven development establishes the specification as the authoritative source of truth and treats code as a derivative artefact. It replaces loose, non-deterministic prompting with unambiguous, executable contracts grounded in architectural and security constraints. That supports the traceability the Australian Information Security Manual and ISO 27001 require, validating every change against documented requirements through automated gates in CI/CD. Security corrections propagate across future regeneration cycles, which mitigates architectural drift and preserves individual human accountability for AI-enabled outcomes.

The three-phase lifecycle

Phase 1
Spec Writing
Interface contracts, data shapes, acceptance criteria, file layout, security constraints, test scope. Quality here has compounding leverage.
8–15 turns · 15–30 min
Opus 4.8, justified for all tiers
~4,000 tok regular input / turn
Output: 1,500–2,500 token spec
Phase 2
Spec Execution
The spec replaces file-discovery turns. Start a fresh session; the spec lives in a file.
40–65 turns · up to 4 hrs
Sonnet 5 primary, Haiku 4.5 for simple turns
~13,000 tok regular input / turn
/compact at 80% context fill
Phase 3
Conformance Review
Pattern matching against spec criteria. Failures return to Phase 2 with specific deviation notes.
5–10 turns · 15–30 min
Haiku 4.5 primary, Sonnet 5 for edge cases
~3,000 tok regular input / turn
Starts fresh, references the spec file

Why it works: the front-loading principle

The dominant agentic cost driver is regular input from file exploration and growing conversation history. In a standard agentic session Claude reads ten to fifteen files to discover scope before work begins, and each read compounds into every subsequent turn. A spec replaces that discovery phase with a single authored document. Claude reads the spec and the specific files it names, and nothing else.

25,000Average regular input tokens per turn, standard agentic heavy
~13,000Average regular input tokens per turn, spec-driven execution
−48%Reduction in the dominant cost driver

The second saving is wrong-direction turns. A standard agentic session at heavy usage accumulates four to eight turns building something later discarded because a constraint was not understood upfront. A spec surfaces those errors at Phase 1 review, where correction costs roughly 500 tokens rather than tens of thousands.

Session sequence: clear between Phase 1 and Phase 2

Write the specification to a file before the authoring session ends, then start a fresh session to execute it. The execution session then carries clean context focused entirely on implementation and references the written spec. That sequence is also cheaper, because a fresh session carrying a 2,000-token spec file costs less than one carrying the full authoring history.

✗ Carrying the authoring history
# Phase 1 complete, spec written # Continue Phase 2 in the same session Now execute against the spec...
Every Phase 2 turn carries the full authoring history, including the alternatives you rejected. That history costs tokens on every turn and distracts Claude during implementation.
✓ Fresh execution session
# Phase 1 complete Write the spec to SPEC.md before we finish. /rename auth-service-spec /clear # Fresh Phase 2 session Implement @SPEC.md section 3, the AuthService interface. Start with src/auth/service.py.
A fresh session carrying a 2,000-token spec file costs less than one carrying the authoring history, and performs better because the context holds only what implementation needs.
BoundaryRule
Phase 1 → Phase 2Write the spec to a file, end the session, start fresh referencing the spec path
Within Phase 2/compact at 80% context fill. Focus it: /compact focus on spec compliance and decisions made so far
Phase 2 → Phase 3Start fresh, referencing the spec file by path
Between sprints/rename then /clear. A new sprint means a new Phase 1

Let Claude interview you to produce the spec

The most efficient route to a spec is to have Claude ask you the questions. Start Phase 1 with a minimal prompt and let the interview surface what you have not considered, which is precisely the wrong-direction cost the spec exists to remove.

# Phase 1, opening prompt I want to build [brief description]. Interview me in detail using the AskUserQuestion tool. Ask about technical implementation, edge cases, concerns and tradeoffs. Don't ask obvious questions, dig into the hard parts I might not have considered. Keep interviewing until we've covered everything, then write a complete spec to SPEC.md.

What a good spec contains

A spec is not a project overview. It is a set of actionable constraints Claude can follow and verify against. Every line should be something Claude does differently because it is there. The most useful specs are self-contained: they name the files and interfaces involved, state what sits out of scope, and end with an end-to-end verification step that proves the feature works.

○ Not a spec, too vague
Build a notification service. It should support email and SMS and be async. Handle errors properly and write tests.
No interface contract. No queue technology. No error-handling pattern. No coverage requirement. Claude makes assumptions, some wrong, and Phase 2 diverges from intent.
✓ A spec, actionable constraints
Service: NotificationService Interface: send(notification: Notification) -> Result[str, NotificationError] Queue: SQS via @infrastructure/queues.py. Not Redis Retry: circuit-breaker in @src/resilience/circuit_breaker.py max 3 retries, exponential backoff Channels: EmailProvider, SMSProvider both implement ChannelProtocol Errors: NotificationError(channel, reason, retryable: bool) Out of scope: push notifications, delivery receipts, template management Tests: unit per provider + integration per channel, coverage >= 90% File layout: src/notifications/service.py src/notifications/providers/{email,sms}.py src/notifications/errors.py tests/notifications/ Verification: pytest tests/notifications/ passes and scripts/send_test_notification.py delivers to the local SQS stub
Interface contract, queue technology, retry pattern, error type, exclusions, test requirements, file layout and an end-to-end check. Phase 2 proceeds from turn 1 without discovery.

Where the spec lives

The spec belongs in a separate file referenced at session start, never embedded in CLAUDE.md. CLAUDE.md loads on every API call in every session, so a 600-line spec there adds approximately 9,000 tokens of cache-read cost per call with no benefit, including on calls with nothing to do with the sprint. Keep CLAUDE.md under 200 lines with a single reference line: Current sprint spec: @SPEC.md.

CLAUDE.md above 3,000 tokens on a spec-driven project indicates this failure is active. Factor 10 covers the fix.

Why Opus 4.8 is justified at Phase 1 for every tier

Model access is normally gated by developer tier. Spec authoring creates an amortisation justification that applies regardless of tier. A single Opus 4.8 spec session producing a tight 1,500 to 2,500 token specification amortises across 40 to 65 Sonnet 5 execution turns, and the overhead recovers within the first execution session through reduced file-exploration turns and eliminated wrong-direction corrections.

The service implements a spec-write session profile that permits Opus 4.8 and enforces Sonnet-only for subsequent sessions in the same sprint. Phase 1 spec sessions need an explicit Opus permit in the service. Opus usage in Phase 2 or Phase 3 is policy drift and the gateway will stop it.

Telemetry signals

✓ Healthy indicators
  • Phase 2 regular input below 15,000 tok per turn
  • Cache hit rate above 80% in Phase 2
  • CLAUDE.md below 3,000 tokens
  • Opus usage confined to Phase 1
  • Phase 3 on Haiku for 70% or more of turns
○ Warning signals
  • Phase 2 regular input above 18,000 tok per turn, so Claude is still file-exploring
  • CLAUDE.md above 3,000 tokens, so the spec is embedded
  • Opus usage in Phase 2 or 3, which is policy drift
  • More than 2 wrong-direction corrections per Phase 2 session, so revisit Phase 1
  • Phase 3 running Sonnet for all turns
Knowledge Check · Factor 04
Scenario: You have finished Phase 1 and written a 2,000-token spec for a new authentication service. A teammate suggests embedding the spec in CLAUDE.md so it is always available, then continuing Phase 2 in the same session so the spec stays warm in context. Are they right?
A. Yes on both counts
B. Right about staying in the same session, wrong about CLAUDE.md
C. Wrong on both counts. The spec belongs in SPEC.md, and you should end the authoring session and start Phase 2 fresh referencing the spec file
D. Right about CLAUDE.md, wrong about the session
Show answer
C is correct. CLAUDE.md loads on every API call in every session, including sessions unrelated to this sprint, so a 2,000-token spec there is pure overhead. And continuing in the same session carries the whole authoring history, including rejected alternatives, into every execution turn. A fresh session carrying the spec file costs less and performs better.
←
Previous
Factor 03 · Verification Targets
Next
Factor 05 · Session Hygiene
→
Wiki / Optimisation / Session Hygiene
Factor 05 of 12

Session Hygiene

Every API call processes the full conversation history to date. Accumulated irrelevant turns add tokens to every subsequent message and degrade output quality, because performance falls as the context window fills.

TL;DR
Run /clear between every unrelated task in DevSecOps work. After correcting Claude twice on the same issue, clear and start fresh with a better prompt. Prefer /clear over /compact, because compaction is itself a large request. Use /btw for questions that do not need to stay in context.
💰
Cost impact: high DevSecOps, important agentic
A 20-turn session where turns 15 to 20 address a different task than turns 1 to 14 carries fourteen turns of irrelevant context as overhead on those last six messages. Each pays full input price for context it will never use.
15 to 25% regular input reduction from consistent hygiene

The commands

CommandWhat it doesWhen
/clearResets the context window completelyBetween unrelated tasks. After two failed corrections. Switching project or codebase
/renameNames the session so you can resume itBefore /clear, whenever you might return
claude --resume
claude --continue
Returns to a named session instead of rebuilding contextWork spanning multiple sittings. Treat named sessions like branches, one per workstream
/compactSummarises history, keeping code and decisionsLong agentic sessions at 70 to 80% context fill, where continuity matters
/compact focus on XFocuses the summaryContinuing with a specific subset of a long session
/btwAsks a side question whose answer never enters conversation historyChecking a detail without growing context. The cheapest question you can ask
Esc Esc or /rewindRewind menu, with summarise-from-here and summarise-up-to-hereCondensing part of the conversation while leaving the rest intact
/contextShows what is currently consuming contextBefore deciding whether to compact or clear
Prefer /clear over /compact. /compact reads the conversation it summarises, so compacting a large context is itself a large request, and the following turn writes a fresh prefix because summarisation replaces the conversation prefix. /clear costs nothing. Reserve /compact for continuity within one task, and where you only need part of the conversation condensed, the rewind menu costs less than a full compaction.

DevSecOps: every task boundary is a clear boundary

✗ No hygiene, session drift
Turn 1: Review @auth/login.py for SQL injection Turn 5: Good. Now check the session middleware too Turn 9: And while we're at it, API rate limiting? Turn 13: Can you also look at the Redis config? Turn 17: One more thing, the password reset flow…
By turn 17 the context holds auth code, session middleware, rate limiting analysis, Redis configuration and password reset logic. Every new message pays for all of it, and Claude's attention divides across five unrelated topics.
✓ Clean boundaries
Turn 1: Review @auth/login.py for SQL injection Turn 5: Looks good, thanks. /rename auth-sql-review-2026-08-26 /clear Turn 1: Review @middleware/session.py for session fixation Turn 4: Done. /clear Turn 1: Check rate limiting in @api/throttle.py…
Each task starts with only the fixed system context plus the immediate task. Every session stays micro-sized, and total cost is a fraction of the drifted session.
The two-correction rule. If you have corrected Claude more than twice on the same issue in one session, the context is cluttered with failed approaches. Run /clear and restate the task with the correction built into the prompt. A clean session with a better prompt almost always outperforms a long session with accumulated corrections.

The 5-minute cache is the default, and it refreshes for free

Prompt caching is the largest cost lever in the service, cutting cost by 38 to 63% against an uncached baseline. It is enabled by default on the Bedrock API and the 5-minute TTL applies unless a request specifies otherwise. Two mechanics matter to how you work.

  • 1The TTL resets on every hit, at no charge. The cache expires only if no hits occur within the window, and each use refreshes it for free. A Claude Code session is a tool-use loop in which one conversational turn issues many API requests seconds apart, so the gaps that matter are genuine idle periods rather than turn boundaries. Measured cache hit rates exceed 90% on the 5-minute default.
  • 2Editing CLAUDE.md invalidates the cache. Where a change is a one-off constraint, state it in the prompt instead and update CLAUDE.md at the end of your working block.
Do not enable ENABLE_PROMPT_CACHING_1H. It is a process-level environment variable, not a per-request setting, so it applies to every session you run and cannot be selected per session type. A 1-hour write costs 2.0 times base input against 1.25 times for a 5-minute write, and that 60% premium applies to every cache-write token rather than only the opening prefix. Enabling it costs a heavy DevSecOps developer $10 to $14 per month and a heavy agentic developer $3 to $4, for a benefit the telemetry does not show.

The narrow cases where a 1-hour profile is warranted

Five situations genuinely produce expiry. Each needs measured evidence before the service enables the profile for your developer account, because it is a per-developer-profile decision rather than a per-session one.

CaseWhy the 5-minute cache expires
Long single turnsThe lifetime measures from the start of the request, not the end of the response, and generation time counts against it. A four-minute stream leaves about one minute for the follow-up. Opus 4.8 at high effort is the exposure, and it is the one interactive case that expires a 5-minute cache without you pausing at all
Human review gapsMerge request review, compliance reading and multi-file approval, particularly in Manual permission mode
CI pipeline waitsGitLab MCP workflows that block on a pipeline
Cross-session cache groupingImpossible on a 5-minute TTL. Bedrock uses organisation-level cache isolation, so sharing is available. Measure the benefit before claiming it
Rate limit headroomCache hits are not deducted against rate limits, so fewer re-writes buys throughput on a constrained quota

The long context effect

Sonnet 5, Opus 4.8 and Sonnet 4.6 carry a 1M token context window at standard rates with no long-context surcharge. A larger window delays auto-compaction, so sessions accumulate more context before summarisation and per-turn regular input rises even though the per-token rate does not. Treat it as a token-volume effect and clear more often rather than relying on auto-compaction to arrive.

Persistent instructions belong in CLAUDE.md

If you restate the same instruction every session, such as "always use our custom logger, never print()", it belongs in CLAUDE.md rather than in your prompt. Instructions in CLAUDE.md survive /clear. Instructions in conversation history do not. The sprint spec, however, is not CLAUDE.md material: see Factor 04.

Knowledge Check · Factor 05
Scenario: You are at turn 8 of a DevSecOps session. Claude has twice used float() for currency amounts when your codebase requires Decimal(). What next?
A. Correct it a third time and hope it sticks
B. Run /clear, then restart with a prompt stating "all currency amounts must use Python's Decimal type, never float(), this is non-negotiable"
C. Run /compact and continue
D. Switch to a different model
Show answer
B is correct. Two corrections on the same issue is the trigger. The context now holds two failed attempts, and a fresh session with the constraint stated upfront is faster, cheaper and more likely to be correct. C keeps the failed approaches in the summary and costs a large request to produce it. Add the rule to CLAUDE.md so you never state it again.
←
Previous
Factor 04 · Spec-Driven Development
Next
Factor 06 · Model Selection & Effort
→
Wiki / Optimisation / Model Selection
Factor 06 of 12

Model Selection & Effort Levels

Defaulting to Sonnet for everything is the most common unnecessary cost. Haiku handles a quarter to a third of DevSecOps work and all spec-driven conformance review. Effort levels tune reasoning depth without changing model.

TL;DR
Escalate in order: effort level, then session profile, then model tier. Haiku 4.5 for routine work and Phase 3 conformance. Sonnet 5 for review, analysis, bug fixes and Phase 2 execution. Opus 4.8 for Phase 1 spec authoring, and for novel reasoning with senior approval. Set effort once per session and hold it.
💰
Cost impact: high DevSecOps, moderate agentic
Haiku 4.5 costs $1.10 per MTok input against Sonnet 5 at $2.20, and roughly 2.6 times less for the same text once the tokenizer difference is included, because Haiku retains the previous-generation tokenizer. In DevSecOps work where 25 to 30% of interactions are genuinely routine, consistent Haiku routing produces meaningful monthly savings.
Haiku target: 25–30% DevSecOps · 10–15% agentic standard · ~20% spec-driven

Escalate in the right order

Three levers answer a hard task, and model tier is the most expensive of them. The gateway enforces this ordering rather than leaving it to judgement.

  • 1stEffort level. Sonnet 5 at high effort is the correct first escalation for a standard developer facing a hard task.
  • 2ndSession profile. Plan mode, or a written spec, removes more wrong-direction cost than a larger model adds.
  • 3rdModel tier. Opus 4.8 is the exception, not the default response to difficulty.
  • 4thMax effort on Opus 4.8. Senior-approved sessions only, one turn at a time. The most expensive request available on this deployment.

Model selection guide

Task typeModelWhy
Docstrings, comments, type annotationsHaiku 4.5Pattern completion, no deep reasoning needed
Pipeline failure triage and CI configurationHaiku 4.5Log pattern matching, not novel analysis
Syntax fixes, formatting, boilerplate scaffoldingHaiku 4.5Mechanical transformation, no design decisions
Dependency version and compatibility lookupsHaiku 4.5Factual retrieval
GitLab issue summarisationHaiku 4.5Text processing
Spec-driven Phase 3 conformance reviewHaiku 4.5Pattern matching against defined criteria, not reasoning
Code review and security analysisSonnet 5Requires understanding of intent and edge cases
Bug fix analysis and implementationSonnet 5Reasoning about cause, effect and constraints
Test generation beyond trivial coverageSonnet 5Understanding behaviour under failure modes
Multi-file feature implementationSonnet 5Sustained multi-turn reasoning
Spec-driven Phase 2 executionSonnet 5Haiku eligible for simple bounded implementation turns
Known-pattern threat modelling, OWASP Top 10, CVE triageSonnet 5Established patterns rather than novel reasoning
Automated tasks with no human in the loopSonnet 5Nobody verifies the reasoning, so take the better cost-reliability trade
Spec-driven Phase 1 spec authoringOpus 4.8, all tiersAmortised across 40 to 65 Sonnet 5 execution turns
Novel exploit chain assessmentOpus 4.8, seniorNovel reasoning under genuine ambiguity
Compliance gap analysis against complex controlsOpus 4.8, seniorMulti-control reasoning, high downstream error cost
Greenfield architectural designOpus 4.8, seniorOne planning session prevents many wrong-direction turns

What the models cost relative to each other

ComparisonRate ratioEffective ratio for the same text
Haiku 4.5 against Sonnet 52.0 timesApproximately 2.6 times, since Haiku uses the previous tokenizer
Opus 4.8 against Sonnet 52.5 times2.5 times, since both use the new tokenizer
Opus 4.8 against Haiku 4.55.0 timesApproximately 6.5 times

Why Sonnet 5 and Opus 4.8 rather than the 4.6 generation

ComparisonWhere the current model is betterWhere it costs more
Sonnet 5 against Sonnet 4.6Anthropic positions it as its most agentic Sonnet. The rate card is 33% lower, and after the tokenizer increase measured spend is 13.3% lower for identical work, because input, output and both cache rates scale by the same factorNothing. Sonnet 5 costs less on the rate card and less in measured spend
Opus 4.8 against Opus 4.6Minimum cache checkpoint falls from 4,096 tokens to 1,024, which widens what can be cached in short sessions. Adaptive reasoning brings effort-level control. A new system instruction can be added partway through a conversation without invalidating the system or message cachesApproximately 30% more expensive in practice at an identical rate card, because of the tokenizer
Sonnet 4.6 and Opus 4.6 exist for validated legacy workloads only. The reason to retain them is migration cost, not model quality: where an application, prompt library, evaluation suite or agent has been built and validated against a specific model, moving it requires re-testing, and any behavioural difference requires remediation before it can be trusted in production. In a high compliance environment that work carries its own evidence burden. If you are working on such a workload, stay on 4.6 and let your team schedule the migration as planned work. For any new workload, select Sonnet 5 or Opus 4.8.

The Haiku caching constraint, in proportion

Haiku 4.5 requires 4,096 tokens per cache checkpoint against 1,024 for Sonnet 5, and AWS evaluates that minimum against the combined tools, system and messages token count rather than each section individually. With 21,000 tokens of fixed context in a normal session, the threshold clears comfortably. It constrains only a subagent invocation carrying a genuinely small prompt and no tool definitions, where nothing caches and no error is returned. Cache token counts at zero are the only signal.

Switching model mid-session

# Start a session in Haiku for triage work /model haiku Tell me why pipeline #4821 test-unit fails, based on this log: [paste log] # Issue identified. Switch to Sonnet 5 to implement the fix /model sonnet Fix the ImportError in @tests/test_cache.py. TokenCache was renamed to CacheToken in the last commit.

Conversation history carries over and new API calls bill at the new model's rate. The service enforces model access by role-based routing, per user, per team and organisation wide, so a model outside your permitted tier is rejected. A rejection is policy working rather than a fault to raise.

Effort levels: four levels, set once per session

Effort controls how deeply Claude reasons, separately from which model you use. The resolved effort value renders into the prompt, so changing it between requests invalidates the message-block cache and forces a re-write. Choose the effort level with the session type at the start, then steer individual turns through prompt wording.

LevelCommandAvailabilityWhen
Low/effort lowAll modelsRoutine work: docstrings, syntax fixes, formatting, quick lookups, conformance review
Medium (default)/effort mediumAll modelsMost coding work: code review, bug fixes, standard implementation, agentic execution
High/effort highAll modelsSecurity analysis, architectural decisions, complex bug investigation, spec authoring
Max/effort maxOpus 4.8 onlyThe most complex reasoning. Senior-approved sessions only, and it does not persist
Max sits outside the set-once rule because it does not persist. Treat it as a single deliberate call on one hard turn rather than a session setting, and expect the cache write that follows when the resolved effort value changes back. Do not reach for it before the escalation ladder above has been worked through.

Match the level to the session type

Session typeEffort
DevSecOps micro and standardLow or medium
DevSecOps extended, agentic executionMedium
Spec-driven Phase 3 conformanceLow
Spec-driven Phase 1 authoring, novel reasoning under ambiguityHigh
Opus 4.8 on a single turn of genuinely novel reasoning, senior approvedMax
Why the Haiku share differs between patterns. In DevSecOps a substantial fraction of daily interactions genuinely do not require Sonnet-level intelligence. In standard agentic sessions even simple tasks involve sustained multi-turn reasoning where Haiku creates quality drag and more wrong-direction attempts. In spec-driven agentic, Phase 3 conformance review is pattern matching against defined criteria and Haiku handles it at a third of the cost.
Knowledge Check · Factor 06
Scenario: You start a Sonnet 5 session at medium effort for a multi-file refactor. Six turns in you hit a genuinely hard architectural question. What do you do?
A. Switch to /effort high for that turn, then back to medium
B. Switch to Opus 4.8
C. Finish the current session, then start a fresh session at high effort for the architectural question, referencing the specific files
D. Keep going at medium and hope
Show answer
C is correct. Changing effort mid-session invalidates the message-block cache and forces a re-write, so A costs more than it looks. B skips two cheaper levers and reaches straight for the most expensive one. A fresh session at high effort scoped to the one question keeps the cache clean and gives Claude a context free of the refactor history.
←
Previous
Factor 05 · Session Hygiene
Next
Factor 07 · Extended Thinking
→
Wiki / Optimisation / Extended Thinking
Factor 07 of 12

Extended Thinking

Thinking tokens bill as output tokens, the most expensive category. Anthropic enables extended thinking by default because it materially improves complex planning and reasoning. On the current generation it is also invisible, so you pay for reasoning you never see.

TL;DR
Effort level is the control on Sonnet 5 and Opus 4.8. Set it once per session to match the session type. Thinking display defaults to omitted on both models, so read the token counts the Bedrock response reports rather than estimating from the transcript.

The control depends on the model generation

ModelControl
Sonnet 5, Opus 4.8Adaptive reasoning only. Effort level is the control, and these models ignore a nonzero MAX_THINKING_TOKENS
Sonnet 4.6, Opus 4.6MAX_THINKING_TOKENS applies in combination with CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1. Relevant only on a validated legacy workload

Since Sonnet 5 and Opus 4.8 are the models you will use for new work, effort level is your only lever. Factor 06 carries the four levels and the session-type table, and the rule to set it once and hold it applies here for the same cache reason.

Max effort is the most expensive request on this deployment. /effort max runs on Opus 4.8 only, does not persist, and produces the largest thinking blocks available. Thinking bills as output at $27.50 per MTok on Opus 4.8, which is 12.5 times the Sonnet 5 input rate. Use it on one turn of a senior-approved session where the reasoning is genuinely novel, and never as a standing setting.

Thinking is invisible in the transcript

On Sonnet 5 and Opus 4.8, thinking display defaults to omitted rather than summarised. You will not see the reasoning and you still pay for it. Any cost estimate built from returned content under-reports on these models.

Use /cost at the end of a session and read the usage fields the Bedrock response reports. Factor 12 covers the telemetry.
Long thinking turns interact with the cache. The cache lifetime measures from the start of the request rather than the end of the response, and generation time counts against it. A four-minute stream leaves roughly one minute for the follow-up. Opus 4.8 at high or max effort is the one interactive case that can expire a 5-minute cache without you pausing at all.

Sonnet 5 at high effort against Opus 4.8

Sonnet 5 at high effort, when
  • The problem is complex but well defined
  • You can specify the constraints clearly
  • Cost matters
  • This is your first escalation, always
Opus 4.8 at medium effort, when
  • The reasoning is genuinely novel
  • Security threat modelling with ambiguous attack surfaces
  • Compliance analysis where multiple control interpretations exist
  • Sonnet 5 has already failed twice at high effort
Knowledge Check · Factor 07
Scenario: You set MAX_THINKING_TOKENS=2000 on a Sonnet 5 session to cap reasoning cost across a batch of routine tasks. What happens?
A. Thinking is capped at 2,000 tokens
B. Nothing. Sonnet 5 runs adaptive reasoning and ignores a nonzero value. Set the session to low effort instead
C. Extended thinking is disabled entirely
D. The session fails to start
Show answer
B is correct. The variable applies only to Sonnet 4.6 and Opus 4.6, and only with adaptive thinking explicitly disabled. On the current generation, effort level is the control, and a batch of routine tasks is a low-effort session.
←
Previous
Factor 06 · Model Selection & Effort
Next
Factor 08 · Subagents
→
Wiki / Optimisation / Subagents
Factor 08 of 12

Subagents for Research and Cost Routing

Subagents run in separate context windows and report back summaries, so verbose exploration never enters your main conversation. They also serve as cost routers, letting a Sonnet 5 session buy Haiku-priced investigation.

TL;DR
Ask Claude directly: "use subagents to investigate X." Specify model: haiku for file scanning, documentation lookup and log analysis. Subagents are enabled by default and need no configuration. A verification subagent reviewing the diff in a fresh context is the cheapest quality gate available.
💰
Cost impact: high, agentic
Without a subagent, all investigation output stays in the main session and gets re-sent on every subsequent API call. A subagent handles the investigation and returns only its summary, so main session input stays clean.
20 to 40% main session regular input reduction on investigation-heavy tasks

How to invoke one

# Research, routed to Haiku Use a Haiku subagent to check whether redis-py 4.5.4 is compatible with Redis 7.2. Return the compatibility verdict and any breaking changes, nothing else. # Verification, routed to Haiku Use a subagent to verify that all test files in @tests/ import from the correct module paths after the refactor. Return a list of files with broken imports. # Complex investigation, let Claude choose the model Use subagents to investigate why the auth middleware adds 200ms latency. Check the chain in @src/middleware/ and return the most likely cause.

What to route and what to keep

✓ Route to a subagent
  • Library and version compatibility checks
  • Scanning for a pattern across many files
  • Verifying test coverage for a module
  • Checking whether a dependency is present
  • Summarising a large file or log
  • Running shell commands and returning results
○ Keep in the main session
  • Work needing the full conversation context
  • Implementation decisions depending on prior turns
  • Security analysis requiring nuanced interpretation
  • Anything where you need to review Claude's reasoning
  • Anything modifying files in your project
The one caching case to watch. A subagent invocation carrying a genuinely small prompt and no tool definitions can fall below the minimum cacheable size, in which case nothing caches and no error is returned. AWS evaluates the minimum against the combined tools, system and messages token count, so a normal session clears it easily and this affects only very small subagent calls. Cache token counts at zero are the only signal. See Factor 06.

The adversarial review subagent

A verification subagent reviewing the diff in a fresh context is the cheapest quality gate available, because it sees only the diff and the criteria rather than the reasoning that produced the change. The agent doing the work is not the one grading it.

Run the bundled /code-review skill for a correctness check on the current diff. To check the diff against your plan or spec instead, write the prompt yourself.

Use a subagent to review the rate limiter diff against SPEC.md. Check that every requirement is implemented, the listed edge cases have tests, and nothing outside the task's scope changed. Report gaps that affect correctness or the stated requirements, not style preferences.
That last clause matters. A reviewer asked to find gaps will report some even when the work is sound, because that is what you asked for. Chasing every finding produces over-engineering: extra abstraction layers, defensive code, and tests for cases that cannot happen.

Because the reviewer runs as a subagent, the implementing session receives the gaps directly and can fix and re-review without you copying findings between windows.

Subagents against agent teams

Single agentSubagentAgent team
Context windows12, main plus subagent1 per member
Token multiplier1 times~1.2 to 1.5 times~7 times per member in plan mode
Enabled by defaultYesYesNo
Approval neededNoNoYes

If you are unsure which you need, use a subagent. Cost Governance covers agent teams.

Knowledge Check · Factor 08
Scenario: You need to know whether three third-party libraries appear in requirements.txt and at what versions. Cheapest approach?
A. Ask Claude directly, it will read the file
B. "Use a Haiku subagent to check @requirements.txt for flask, celery and redis-py. Return the exact version pins or 'not present' for each"
C. Read the file yourself and paste the three lines
D. Use plan mode to investigate the dependencies
Show answer
B is correct, with C close behind. B reads the file in a separate context on the cheapest model and returns three lines. C also keeps the main context clean and costs nothing at all. A pulls the whole file into main session context, where it stays for every subsequent turn. D misuses plan mode, which exists for code changes rather than lookups.
←
Previous
Factor 07 · Extended Thinking
Next
Factor 09 · Preprocessing Hooks
→
Wiki / Optimisation / Preprocessing Hooks
Factor 09 of 12

Preprocessing Hooks

A hook preprocesses data before Claude sees it. Instead of Claude reading a 10,000-line log to find errors, a hook greps for what matters and returns only the matching lines, cutting context from tens of thousands of tokens to hundreds.

TL;DR
For a brownfield estate with verbose CI output, this is the largest untapped lever available to DevSecOps developers. Hooks are deterministic where CLAUDE.md instructions are advisory, so they also suit any control that must happen every time. Claude will write hooks for you.
💰
Cost impact: high, DevSecOps
A single large log read consumes tens of thousands of tokens and stays in context for every subsequent turn. A PreToolUse hook reduces the same read to hundreds of tokens, every time, without you remembering to ask.
Tens of thousands of tokens down to hundreds, on every read

The mechanism

✗ No hook
Read @ci/logs/pipeline-4821.log and tell me why the test-unit job failed. → Claude reads 40,000 lines → ~55,000 tokens enter context → Every subsequent turn pays for all of it → You do this three times a week
Most of a day's DevSecOps budget on one read, and the budget goes again on the next triage.
✓ PreToolUse hook
Read @ci/logs/pipeline-4821.log and tell me why the test-unit job failed. → Hook greps ERROR|FAIL|Traceback → ~600 tokens enter context → Claude sees only the failure lines → Happens automatically, every time
The reduction is deterministic. It does not depend on you remembering to scope the read.

Hooks are guarantees, CLAUDE.md is a request

A CLAUDE.md line saying "never modify migrations" is advisory, and a long CLAUDE.md means Claude may not read it at all. A hook blocking writes to migrations/ is deterministic. Anywhere a control must happen every time with zero exceptions, use a hook rather than an instruction.

Claude will write them for you

# Ask in plain language Write a hook that runs ruff after every file edit and returns only the violations. Write a PreToolUse hook that truncates any file read over 2,000 lines to the first and last 200 lines with a note. Write a hook that greps pipeline logs for ERROR and WARN and returns only those lines. Write a hook that blocks writes to the migrations folder.

Run /hooks to browse what is configured, and edit .claude/settings.json directly to configure by hand.

On this deployment, hooks are restricted to managed sources. A ConfigChange hook logs in-session settings changes. Raise a new hook through your team's configuration path rather than adding it locally, and expect it to be reviewed like any other managed control.

Where hooks pay most

SituationHook
Verbose CI and pipeline outputGrep for error and failure patterns before the read reaches context
Large generated files that Claude occasionally readsTruncate above a line threshold with a note explaining the truncation
A lint or type check that must run after every editPostToolUse hook returning only violations
A directory that must never be writtenPreToolUse block, which is stronger than a CLAUDE.md rule
An unattended run that must not finish until tests passStop hook, per Factor 03. Budget up to 8 extra turns
Knowledge Check · Factor 09
Scenario: Your CI job produces a 40,000-line log and you triage failures from it several times a week. Claude reads the whole thing each time. Best fix?
A. Ask Claude to read only the last 500 lines
B. Switch to Haiku for triage so the reads cost less
C. A PreToolUse hook that greps for ERROR and returns only matching lines, so the reduction happens every time without you asking
D. Paste the log into the prompt yourself
Show answer
C is correct. A and D work once and depend on you remembering. B reduces the rate but still pays for 40,000 lines of input on every triage, and Haiku is the wrong model to hand a wall of unfiltered noise. The hook is deterministic and applies to every session on the project, including sessions run by your teammates.
←
Previous
Factor 08 · Subagents
Next
Factor 10 · CLAUDE.md & Context
→
Wiki / Optimisation / CLAUDE.md & Context
Factor 10 of 12

CLAUDE.md, Skills & Context Discipline

CLAUDE.md loads into every session automatically. Every token in it is billed at cache-read price on every API call in every session on that project. Keep it sharp: instructions Claude follows, not background Claude reads.

TL;DR
Keep CLAUDE.md under 200 lines and fully actionable. Move anything only sometimes relevant into skills, which load on demand. Configure your ignore rules before the first session on any brownfield project. Connect the GitLab MCP server only for sessions that need it.
💰
Cost impact: medium, both patterns
A bloated 600-line CLAUDE.md with explanatory prose adds roughly 9,000 tokens of cache-read cost per call and produces no better output than a focused 150-line file. It also causes Claude to ignore the instructions that matter, so the discipline serves quality as well as cost.
Target: under 200 lines · 100% actionable · zero narrative prose

What belongs in CLAUDE.md

Run /init to generate a starter file from your project structure, then prune it. For each line, ask whether removing it would cause Claude to make a mistake. If not, cut it. Run /doctor on a checked-in file and Claude proposes cuts for content it can derive from the codebase.

✓ Include
  • Bash commands Claude cannot guess
  • Code style rules that differ from defaults
  • Security rules, such as never logging PII
  • Test patterns and required coverage
  • Key file paths for common reference
  • Forbidden patterns Claude should refuse
  • Developer environment quirks
○ Exclude
  • Anything Claude can work out by reading the code
  • Standard language conventions Claude already knows
  • Detailed API documentation, link to it instead
  • Why you chose a technology
  • Architecture evolution narrative
  • Team onboarding context
  • Information that changes frequently
Emphasis works once. If Claude keeps skipping one instruction, add IMPORTANT to that line alone. If you emphasise many lines, none of them stands out. Check CLAUDE.md into git, treat it like code, and prune it when things go wrong.

CLAUDE.md template, secd3v projects

# Project: [service-name] # Environment: ap-southeast-2 · PROTECTED ## Stack Python 3.12 · Django 5.1 · PostgreSQL 16 · Redis 7.2 · Celery 5.4 AWS: ECS Fargate, RDS, ElastiCache, SQS, S3 ## Critical rules, follow always - All currency: use Decimal, never float - All logging: use src/logging/structured.py, never print() - All API responses: use ResponseWrapper from src/api/base.py - No raw SQL string interpolation, parameterised queries only - Never log: passwords, tokens, PII, credit card data - Migrations: never auto-generate, write manually and review ## Test requirements pytest · fixtures in tests/conftest.py Coverage must not drop below 85% Security-related functions: always add failure-mode tests ## Key paths src/api/base.py base views and ResponseWrapper src/auth/ all authentication logic src/logging/ structured logging utilities infrastructure/ IaC, do not modify in Claude sessions ## Forbidden Do not modify: migrations/, infrastructure/, .env files Do not use: requests (use httpx), time.sleep() (use asyncio) ## Current sprint spec @SPEC.md

Skills for anything only sometimes relevant

Domain knowledge and workflows that apply to some sessions belong in skills, not CLAUDE.md. Claude loads them on demand without bloating every conversation.

# .claude/skills/api-conventions/SKILL.md --- name: api-conventions description: REST API design conventions for our services --- # API Conventions - Use kebab-case for URL paths - Use camelCase for JSON properties - Always include pagination for list endpoints - Version APIs in the URL path (/v1/, /v2/)

Skills also define repeatable workflows you invoke directly with /skill-name. Use disable-model-invocation: true for workflows with side effects you want to trigger manually.

The spec-driven bloat failure

✗ Spec embedded
# CLAUDE.md, 350 lines ## Project rules… ## Current sprint spec: Service: NotificationService Interface: send(notification: Notification)… [300 more lines of spec content]
Every API call in every session pays cache-read cost for 2,000+ tokens of spec, including Phase 3 review, unrelated DevSecOps sessions and CI automation.
✓ Spec externalised
# CLAUDE.md, 45 lines ## Project rules… ## Current sprint spec: @SPEC.md # SPEC.md, separate file, 300 lines Service: NotificationService Interface: send(notification: Notification)…
CLAUDE.md stays under 200 lines. SPEC.md loads on demand during Phase 2. Sessions that do not need the spec do not pay for it.
Telemetry signal. CLAUDE.md above 3,000 tokens on a spec-driven project means the spec has been embedded. Run /context, check the CLAUDE.md line, strip it to SPEC.md and leave a single reference line.

.claudeignore on day one

Configure your ignore rules before the first session on any brownfield project. A single accidental glob-all read of a large repository consumes 50,000 to 150,000 tokens in one call, which is most of a day's DevSecOps budget. Years of accumulated build artefacts, test output, generated files and cached data easily total millions of tokens.

Configure .claudeignore on day one. Cover build outputs, dependency directories, generated files, binaries, large media and IDE directories.

MCP context and connection hygiene

The GitLab MCP server adds 9,000 tokens of tool names and schemas to the fixed cached context, and they load upfront on every session where the server is connected.

ComponentNo MCPWith GitLab MCP
System prompt and agent instructions3,900 tok3,900 tok
Built-in tool definitions15,600 tok15,600 tok
CLAUDE.md and memory files1,500 tok1,500 tok
GitLab MCP tool names and schemasnot applicable9,000 tok
Total fixed cached context21,000 tok30,000 tok
Deferred tool loading does not apply on this deployment. Claude Code can defer MCP tool definitions and load them on demand through tool search, which would cut the GitLab overhead to tool names alone. That saving does not reach here, for two independent reasons. First, the deferral threshold is never reached: deferral engages only when MCP tool descriptions exceed roughly 10% of the context budget, and one GitLab server at 9,000 tokens sits far below that on a 200,000-token window and further below it on a 1M-token window. Second, tool search is not reliable on the Bedrock API: Anthropic excludes the tool search beta header for Bedrock, and an open Claude Code defect records tool search activating client-side and generating tool reference blocks that Bedrock rejects with a 400 error. Plan on 30,000 tokens of fixed context whenever GitLab MCP is connected.

The invalidation risk here comes from the managed MCP configuration changing between sessions, not from loading a tool within one. Cache checkpoints chain in the order tools, then system, then messages, so a change to the tools section invalidates the system and messages caches with it.

ScenarioMCP configuration
Pure coding session, no GitLab operations neededDisconnect. You are paying 9,000 tokens of cached context per session for nothing
Merge request review, reading diffs, checking pipelinesConnect. The overhead earns its place
Mixed session with occasional GitLab checksConnect, and consider whether a CLI covers the occasional check instead. See Factor 11
A server you expected is missingCheck claude mcp list. managed-mcp.json lists only the CCS GitLab MCP server, so nothing else will connect

GitLab calls also add per-turn cost beyond the fixed context

A GitLab call returns a payload that enters the conversation as an uncached tool result, whether that is a merge request diff, a pipeline log, an issue body or a review comment thread. That sits on top of the 9,000-token difference in fixed cached context.

PatternAdditional regular input per turn EstWhy
DevSecOps~1,010 to 1,122 tokMerge request review and pipeline triage call GitLab on most turns
Agentic~450 to 700 tokA long autonomous session touches GitLab less often relative to its turn count

Output rises by up to 105 tokens per turn as well. These uplifts are the second-largest input to the with-MCP cost figures after the fixed context.

Scope large diffs. A merge request diff on a big feature branch returns 30,000 to 50,000 tokens, all of it uncached and all of it staying in context for the rest of the session. Ask for less: "Get MR !847's diff but only read and summarise the changes in src/auth/, skip other directories."
Knowledge Check · Factor 10
Scenario: Your CLAUDE.md holds a 400-word explanation of why the team chose Django over FastAPI, a description of the branching strategy, an architecture history, and the full 2,000-token sprint spec. Your team is hitting high token costs. What do you fix first?
A. Keep it all, context helps Claude make better decisions
B. Remove the Django rationale and branching strategy only
C. Extract the spec to SPEC.md first, since it costs on every call in every session. Then cut the rationale, branching strategy and architecture history. Keep architecture decisions only where they read as rules Claude can follow
D. Compress everything to 200 words and keep it in CLAUDE.md
Show answer
C is correct. The spec is the highest-impact fix because it adds cost to sessions that have nothing to do with the sprint. The rationale, branching strategy and history are not instructions Claude follows differently on any specific task. "Always use ResponseWrapper" belongs. "We moved to Django in 2024 because" does not.
←
Previous
Factor 09 · Preprocessing Hooks
Next
Factor 11 · CLI Tools & Plugins
→
Wiki / Optimisation / CLI Tools & Plugins
Factor 11 of 12

CLI Tools & Code Intelligence Plugins

CLI tools are the most context-efficient way to reach an external service, because they add no per-tool listing to your context window. A code intelligence plugin replaces text search with precise symbol navigation.

TL;DR
Where both a CLI and an MCP server exist for the same service, prefer the CLI, which matters more here because the GitLab MCP schemas load upfront rather than on demand. Claude learns tools it does not already know from --help. For typed languages, install a code intelligence plugin so one go-to-definition call replaces a grep and several speculative file reads.
💰
Cost impact: medium, both patterns
A CLI adds nothing to the context window at all. This matters more on this deployment than it would elsewhere, because the GitLab MCP schemas load upfront rather than on demand, so an MCP server you connect costs 9,000 tokens of cached context whether or not you call it.
Zero context overhead against 9,000 tokens for the GitLab MCP server

Teach Claude a tool it does not know

Use 'foo-cli-tool --help' to learn about foo, then use it to solve A, B and C.

Claude is effective at learning CLI tools from their own help output, so an unfamiliar internal tool is not a reason to reach for an MCP server.

Code intelligence for typed languages

✗ Text search
grep -r "TokenCache" src/ → 14 matches across 9 files → Claude reads 5 candidate files → ~12,000 tokens to find one definition
Each speculative read stays in context for the rest of the session.
✓ Code intelligence plugin
go-to-definition: TokenCache → src/cache/token.py:42 → ~200 tokens Language server also reports type errors automatically after each edit.
One call replaces a grep and several reads, and Claude catches type mistakes without running a compiler.

Run /plugin to browse the marketplace. Plugins bundle skills, hooks, subagents and MCP servers into a single installable unit.

On this deployment, plugin and CLI tool additions go through your team's approved software path. The managed-mcp.json position in Factor 10 exists for the same reason: an MCP server is remote code with tool authority over the agent, and it gets governed like software you install.
Knowledge Check · Factor 11
Scenario: You need Claude to read issue details, open merge requests and check pipeline status on a project where both the GitLab MCP server and a GitLab CLI are available. Which costs less in context?
A. The MCP server, because it is purpose-built for the task
B. The CLI, because it adds no per-tool listing to the context window
C. No difference, both make the same API calls
D. Whichever loads first in the session
Show answer
B is correct. The GitLab MCP server loads 9,000 tokens of tool names and schemas upfront on every session it is connected to, and deferred loading does not apply here. A CLI adds nothing to the context window. The MCP server still earns its place where it enforces a security posture the CLI cannot, which is why the CCS GitLab MCP server is the managed option and its default posture suppresses mutating tools from the surface entirely.
←
Previous
Factor 10 · CLAUDE.md & Context
Next
Factor 12 · Token Telemetry
→
Wiki / Optimisation / Token Telemetry
Factor 12 of 12

Token Telemetry

Without instrumentation, optimisation is guesswork. Run /context at the start of any session where cost looks wrong, and /cost at the end of any session to see where the money went.

TL;DR
If your cache hit rate sits under 70%, your DevSecOps regular input exceeds 15,000 tokens per turn, or your CLAUDE.md exceeds 3,000 tokens on a spec-driven project, something is wrong and fixable. Diagnose whether a cache miss came from expiry or invalidation before changing any setting.

Health metrics

SignalIndicatesWhat to do
Cache hit rate below 70%Poor session hygiene, or one of the invalidation causes belowDiagnose before changing any TTL setting. Below 50%, act immediately
Regular input above 15,000 tok per turn, DevSecOpsBroad prompting, or a missing /clearFactor 01 and Factor 05
Regular input above 18,000 tok per turn, spec-driven Phase 2The spec is not being referenced and Claude is still file-exploringTighten the spec scope, per Factor 04
CLAUDE.md above 3,000 tokens on a spec-driven projectSpec prose embedded in CLAUDE.mdExtract to SPEC.md. This is an endpoint-side check only, since the gateway records metadata and cannot inspect the file
Haiku share below 15% on DevSecOpsModel discipline not appliedFactor 06
Cache write records reporting a 1-hour TTLThe 1-hour setting is active on a profile that may not warrant itRaise it with your platform owner
Cache token counts at zeroA prompt below the minimum cacheable length for the model in useAlmost always a very small subagent call, per Factor 06

Diagnose a cache miss before you change anything

A longer TTL fixes expiry. It does nothing for invalidation and costs the premium anyway. Attribute the miss first.

EvidenceCause
Full-prefix cache write after an idle gapTTL expiry. The only evidence that justifies enabling the 1-hour profile for your account
Full-prefix cache write with no preceding gapInvalidation. Work through the six causes below. The remedy is configuration discipline rather than a longer TTL
Cache write volume rising with no change in session patternRouting behaviour under load where a multi-region profile is in use

The six invalidation causes

CauseEffectRemedy
Effort level changed mid-sessionThe resolved effort value renders into the prompt, so a change invalidates message blocksSet effort once per session and hold it, per Factor 06
Thinking configuration changedSame mechanism as effortSet once per session
Tools section changedCheckpoints chain in the order tools, system, messages, so modifying tools invalidates the system and messages caches with itHold the managed MCP configuration stable between sessions
Image addedAdding an image anywhere in the prompt invalidates message blocksExpect a full re-write after a pasted screenshot
20-block lookback exceededAutomatic prefix checking looks back approximately 20 content blocks from the checkpoint, and static content beyond that range is not foundOccurs on long parallel tool sequences. Additional checkpoints are the documented remedy, up to the four-checkpoint maximum
Cross-region routing under loadAt times of high demand, cross-region inference optimisations may lead to increased cache writesRaise a sustained rise with your platform owner. Direct in-region routing is preferable for cache-sensitive workloads where capacity allows
/compact also produces a full cache write on the following turn, because summarisation replaces the conversation prefix. That is a cost of compaction rather than a cache fault.

Reading a /context breakdown

The fixed cached context is approximately 21,000 tokens without MCP and 30,000 with the GitLab MCP server connected. Factor 10 carries the component breakdown.

Signal in /contextMeaning
CLAUDE.md well above 1,500 tokensBloated. Review and trim per Factor 10
CLAUDE.md above 3,000 tokens on a spec-driven projectSpec embedded. Extract to SPEC.md
Regular input far above the session-type figureSession drift. Consider /clear
GitLab MCP schemas present when you have no GitLab work9,000 tokens of cached context you are paying for and not using. Disconnect
SPEC.md absent during Phase 2Phase 2 running without the spec, so the saving is not being realised

Thinking is invisible, so do not estimate from the transcript

On Sonnet 5 and Opus 4.8, thinking display defaults to omitted. Any cost estimate built from returned content under-reports. Use /cost and read the cache and token counts the Bedrock response reports.

Where the numbers come from

Claude Code calls Bedrock directly and sends no usage metrics to Anthropic, so attribution comes from the request path. Every inference request passes through the CCS gateway, which is therefore the complete point of measurement: it prices and attributes each request, aggregated per user, per team, per model and organisation wide. The gateway also assumes its upstream role per user with session tags, so the AWS audit trail corroborates the gateway record rather than substituting for it.

Your organisation should run a recurring 30-day telemetry review against the signals above, restating the modelled inputs behind the cost figures from measured data.
Knowledge Check · Factor 12
Scenario: You run /context at turn 5 of a DevSecOps session: CLAUDE.md 8,200 tokens, regular input 38,000 tokens, cache hit rate 45%. What are the two urgent problems?
A. Upgrade to Opus 4.8 and enable agent teams
B. CLAUDE.md is roughly five times the efficient size, and regular input at turn 5 is more than twice the DevSecOps target, which means the session has already drifted. Trim CLAUDE.md, then /clear and restart with a single-task prompt
C. Enable the 1-hour cache TTL to fix the hit rate
D. Switch to max effort to compensate for the poor context
Show answer
B is correct. CLAUDE.md at 8,200 tokens is cached prose and history rather than actionable instructions. Regular input at 38,000 tokens by turn 5 means several unrelated topics sit in context, and the 45% hit rate confirms it. C treats a hygiene problem as an expiry problem and pays a 60% write premium for nothing. Neither problem improves by changing model or effort, and D would make it considerably worse.
←
Previous
Factor 11 · CLI Tools & Plugins
Next
Cost Governance
→
Wiki / Governance
Cost Governance

Cost Governance

Two capabilities multiply cost rather than reduce it, and both need approval before use. Six further behaviours consume tokens outside the session shapes this guide describes, and three of them bill a full-context turn while you are doing nothing.

Agent teams

🚨
Cost impact: critical. Requires explicit approval
An agent team is a set of Claude instances working in parallel, each with its own context window. A team uses approximately 7 times more tokens than a standard session when its agent teammates run in plan mode. A heavy agentic developer at $151 per month moves well past $1,000 running teams daily.
~7 times multiplier · disabled by default · service budget ceiling is the primary control
Single agentSubagentAgent team
Context windows12, main plus subagent1 per member
Token multiplier1 times~1.2 to 1.5 times~7 times per member in plan mode
Enabled by defaultYesYesNo, gated behind CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
Approval neededNoNoYes, granted to your engineering team, with a service budget ceiling
SuitsAll workResearch, verification, cost routingGenuinely parallel independent tasks only

Running an agent team: four controls

These apply once your engineering team has been granted approval and you are about to spawn an agent team. Every "teammate" below is a Claude instance, not a colleague.

  • 1Run agent teammates on Sonnet 5. Never Opus 4.8. The multiplier applies to whatever rate you set, so the model choice compounds across every member.
  • 2Keep the agent team small. Each agent teammate is a full multiplier rather than a marginal addition, so a fourth member costs what the first one did.
  • 3Write focused spawn prompts. Agent teammates load CLAUDE.md, MCP servers and skills automatically, so a broad spawn prompt pays for that context on every member.
  • 4Shut agent teammates down when their work is done. An idle agent teammate keeps consuming tokens. Leaving a three-member team running overnight is a significant unintended cost event.
✓ Genuinely parallel
  • Three independent microservices simultaneously
  • Test suites for separate unrelated modules
  • Documentation across disconnected packages
  • Migration scripts for independent database tables
  • Any work with no shared state
○ Use a single session instead
  • Sequential tasks disguised as parallel ones
  • Tasks sharing files or modules
  • Work where output of A feeds input of B
  • Feature implementation, use plan mode and Sonnet 5
  • Any task a single focused session could handle

/batch

/batch splits a change across 5 to 30 subagents, each in its own worktree, each opening a merge request. Treat it as an uncapped multiplier alongside agent teams, governed at the service budget rather than by guidance. Test the instruction on two or three files before running it at scale.

Background token consumers

Six behaviours consume tokens outside the session shapes in The Three Work Patterns, and three of them bill a full-context turn while you are doing nothing.

DriverMechanismControl
Compaction/compact reads the conversation it summarises, so compacting a large context is itself a large request, and the following turn writes a fresh prefixPrefer /clear between unrelated tasks. Reserve /compact for continuity within one task
/goal conditionsA separate evaluator re-checks the goal after every turn, and idle check-ins start a new turn carrying full context while background work runsCLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 where goals are not required
Stop hooksBlocks the turn from ending until the check passes, for up to 8 consecutive blocksKeep the check cheap and fast. Budget up to 8 extra turns per gated task
Scheduled tasksFires on its interval and sends full context whether or not the session is activeDisabled by policy unless a named use case justifies it
Cross-session messagingA message from another session arrives as a new turn with full contextcrossSessionInbound set to hold
Agent teams and /batchMultipliers rather than overheadsApproval and a service budget ceiling

Settings you should not change

Four service settings carry the estate's financial risk. They are set for you, and knowing why helps you recognise a misconfiguration on your own endpoint.

SettingRequirementConsequence if omitted
ENABLE_PROMPT_CACHING_1HUnset, so the Bedrock 5-minute default applies. Enabled per developer profile only on measured expiryEnabling it estate-wide costs $3 to $14 per developer per month for no measured benefit
CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0Set unless goal-driven sessions are an approved workflowIdle check-ins send full context on an interval
crossSessionInbound = holdSetUncontrolled turns on idle sessions
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMSUnset unless your engineering team holds approval, with a service budget ceilingApproximately 7 times token usage
managed-mcp.jsonLists only the CCS GitLab MCP server, and held stable between sessionsUncontrolled MCP context and tool authority, plus cache invalidation when the tools section changes
defaultModePinned to auto mode through managed settings, with the organisational risky action set held as ask rulesSessions fall back to Manual mode, which raises turn counts and monthly cost for identical work and reintroduces approval fatigue on trivia

Overall, the agent teams flag carries the largest single financial risk at roughly 7 times token usage, and the caching default carries the largest recurring one.

←
Previous
Factor 12 · Token Telemetry
Next
Model & Planning Reference
→
Wiki / Reference
Reference

Model & Planning Reference

A lookup rather than a read. Which models you have, how each one caches, why token counts differ between generations, and the planning allowance for work these figures exclude. Rates, monthly cost tables and the model access splits live in the companion cost model.

The models provided

Five models are available through the service, and the gateway enforces which of them your role can reach. Sonnet 5 and Opus 4.8 are the current generation and the right choice for any new work. Sonnet 4.6 and Opus 4.6 exist for workloads already built and validated against them.

ModelRoleMin tokens per checkpointMax checkpointsSupported TTLTokenizer
Haiku 4.5Routine tasks, subagents, spec-driven conformance review4,09645 minutes, 1 hourPrevious generation
Sonnet 5Primary workhorse. Code review, analysis, bug fixes, spec execution1,02445 minutes, 1 hourCurrent
Sonnet 4.6Validated legacy workloads only1,02445 minutes, 1 hourPrevious generation
Opus 4.8Maximum intelligence. Senior-approved, and spec authoring for all tiers1,02445 minutes, 1 hourCurrent
Opus 4.6Validated legacy workloads only4,09645 minutes, 1 hourPrevious generation
The 5-minute TTL is the default and the right choice. All five models support both options, and Factor 05 covers why the 1-hour profile is a narrow exception rather than an upgrade.

How the cache limits behave

Three mechanics govern the limits in the table above, and each has a consequence you can see in your own telemetry.

The minimum is cumulative
The minimum is evaluated against the combined token count across the tools, system and messages sections rather than each section individually. With roughly 21,000 tokens of fixed context in every session, the 4,096-token minimum on Haiku 4.5 and Opus 4.6 clears comfortably. It constrains only a subagent invocation carrying a genuinely small prompt and no tool definitions.
Sections chain
Checkpoints process in the order tools, then system, then messages. Modifying the tools section invalidates the system and messages caches with it, which is why the managed MCP configuration is held stable between sessions.
Inference succeeds silently when caching does not
A request below the minimum still returns a response, with no error, and nothing cached. Cache token counts at zero are the only signal, so it is a telemetry check rather than something you will notice while working.
Four checkpoints is the ceiling
Where a long parallel tool sequence pushes static content beyond the roughly 20-block lookback window, additional checkpoints are the documented remedy, up to that maximum. See Factor 12.

Why token counts differ between generations

The current generation uses a newer tokenizer that produces approximately 30% more tokens for the same text. In this catalogue that separates Sonnet 5 and Opus 4.8 from Sonnet 4.6, Opus 4.6 and Haiku 4.5. Rate cards do not reflect it, so the effect lands entirely on token volume and stays invisible in a price comparison.

ComparisonToken differenceNet effect on spend
Sonnet 5 against Sonnet 4.6approx. +30%Sonnet 5 costs approximately 13% less for identical work, because its lower rate card more than offsets the extra tokens
Opus 4.8 against Opus 4.6approx. +30%Opus 4.8 costs approximately 30% more in practice, because the rate card is identical and only the token count moves
Haiku 4.5unchangedUnchanged. Haiku keeps the previous-generation tokenizer, which is why it is cheaper against Sonnet 5 than the rate cards alone suggest
If your token counts rose after moving to Sonnet 5 while your spend fell, this is why. Report the spend, not the token count. The 30% multiplier is a modelled figure pending confirmation against production telemetry at the 30-day review.

Work these figures exclude, and the 10% allowance

Every cost figure for this service models software development: reading and writing code, reviewing changes, generating tests, auditing for security defects, and building services against a specification. The session shapes in The Three Work Patterns are drawn from that work and nothing else.

You will also use the agent for content activities that sit alongside development and are not costed anywhere: long-form documentation, release notes and changelogs, merge request and release summarisation, architecture and sequence diagrams, commit message drafting, incident write-ups, runbooks, onboarding material, and open-ended explanations of unfamiliar code.

These sessions carry a different token shape. They run short, read a bounded amount of context, and produce a great deal of output, which is the most expensive token category at five times the input rate. That is why a short content session is not free: it is output-heavy where a code session is input-heavy.

One such session per day, modelled at 6 turns with 4,000 tokens of regular input and 1,500 tokens of output per turn on the standard developer split, adds the following proportion to the base figure for each pattern and tier.

Pattern and tier, standard policyOne session / dayTwo sessions / day
DevSecOps light, no MCP16.0%32.1%
DevSecOps heavy, no MCP8.0%15.9%
DevSecOps heavy, GitLab MCP6.3%12.5%
Agentic heavy, no MCP3.4%6.9%
Agentic heavy, GitLab MCP3.2%6.4%
Add 10% to every cost figure as a planning allowance. That corresponds to one to two complementary sessions per developer per day and sits mid-range across the patterns.
UseWhen
10%The default. One to two complementary sessions per developer per day, mid-range across all patterns and tiers
15%Where merge request summarisation, release notes or documentation form a stated part of the role. Also for developers at the light DevSecOps tier, where the allowance is proportionally larger against a smaller base
5%Defensible for a purely agentic team, because a heavy autonomous session dwarfs a short content session

The allowance is a planning figure rather than a measured one. It gets replaced at the 30-day telemetry review by separating sessions whose output-to-input ratio exceeds roughly 1 to 3, which is the signature of content work rather than code work. If your own usage is mostly documentation and summarisation, say so when your team sets budgets, because the base figures will understate you.

Factor quick summary

#FactorImpactKey action
01Prompt specificityHighestFile path, line range, specific concern. 5W1H checklist. One task per message
02Plan modeHighest agenticShift+Tab or claude --permission-mode plan. Skip for one-sentence diffs
03Verification targetsHighGive Claude a check that returns pass or fail. Ask for evidence, not assertion
04Spec-driven developmentHighest agenticAuthor the spec, write it to SPEC.md, end the session, execute fresh. Phase 3 on Haiku
05Session hygieneHigh DevSecOps/clear between unrelated tasks. Prefer /clear over /compact. /btw for side questions
06Model selection and effortHigh DevSecOpsEscalate effort, then session profile, then model. Set effort once per session
07Extended thinkingHigh if unmanagedEffort level is the control on Sonnet 5 and Opus 4.8. Thinking is invisible and still billed
08SubagentsHigh agenticHaiku subagents for research. Adversarial review subagent on the diff
09Preprocessing hooksHigh DevSecOpsCut a 10,000-line log to hundreds of tokens before Claude reads it
10CLAUDE.md, skills and contextMediumUnder 200 lines. Skills for sometimes-relevant knowledge. Spec to SPEC.md, never CLAUDE.md
11CLI tools and pluginsMediumPrefer a CLI over an MCP server. Code intelligence plugin for typed languages
12Token telemetryFoundation/context to diagnose, /cost to attribute. Diagnose invalidation before blaming expiry
←
Previous
Cost Governance
Return to
Overview
↑