Developer Cost
Optimisation Training
Twelve developer behaviours control what Claude Code costs you, your team and your organisation. Each factor states its cost impact, explains the mechanism, gives worked examples from DevSecOps, agentic and spec-driven work, and ends with a knowledge check. Every figure is priced on Amazon Bedrock Australian endpoints with the 10% sovereignty premium applied.
Before the factors
The twelve optimisation factors
What a developer costs
Standard developer policy with no Opus access, per month in USD, 22 working days, Australian endpoint rates. Use these to recognise when your own telemetry has drifted.
Prompt caching is the largest single lever, cutting cost by 38 to 63% against an uncached baseline. The 5-minute default captures that saving with no configuration and no premium. Spec-driven development is the second, cutting heavy agentic spend by 28% for the cost of one authoring session per sprint.
Working Securely
The secd3v Claude Code Service implements ISM PROTECTED controls and can support unclassified through to PROTECTED workloads. Customers can deploy the Claude Code CLI with CCS using a range of different security profiles, for Essential Eight Maturity Level 2 (Guide 1) or ISM PROTECTED (Guide 2). Most of the posture is managed configuration you cannot change. This page covers what depends on you.
Permission modes: use auto mode
Claude Code sessions start in Manual mode, where Claude asks before file writes and most shell commands. Auto mode is opt-in on Amazon Bedrock and must be pinned through managed settings using defaultMode, so it is configured for you rather than something you switch on per session. Anthropic has stated an intention to make it the default on cloud platforms without committing to a date.
Auto mode does not remove the safety gate, it moves it
Instead of prompting on routine work, auto mode routes every tool call through a separate classifier that evaluates whether the action is irreversible, destructive, or aimed outside your environment. Data exfiltration sits in a hard-deny category the classifier is designed never to approve. Actions it judges safe proceed without interrupting you.
Two gates, with different reach
| Gate | Covers | Mechanism |
|---|---|---|
| Agent ask rules | Commands and file operations the agent runs on your endpoint, over the organisationally defined risky set | Managed ask rules, evaluated before the classifier, binding in every permission mode |
| CCS GitLab MCP checkpoints | High-risk platform operations, including merge approvals and mutations on the code platform, which the agent's own controls do not reach | Server-side checkpoints, plus a default posture that suppresses mutating tools from the surface entirely |
The gateway carries no human-in-the-loop mechanism of its own, so these two are the gates. They cover different ground rather than duplicating each other: an ask rule stops a command you run, and a checkpoint stops a platform operation the agent invokes through the MCP server.
--dangerously-skip-permissions is out of scope on this deployment and the service rejects it.
What still asks you, and why it matters
Your ask and deny rules bind in every permission mode and are evaluated before the classifier, so the organisational risky action set below stops and asks by name regardless of what the classifier would have decided. This is the ISM-2113 human approval gate in practice.
| Action | Prompts you | Profile |
|---|---|---|
git push | Always | Guide 1 (E8 ML2) |
| Package publication | Always | Guide 1 (E8 ML2) |
| Deletion outside the working tree | Always | Guide 1 (E8 ML2) |
| Any command carrying a credential | Always | Guide 1 (E8 ML2) |
| Merge approvals | Always | Guide 2 (PROTECTED) |
| Pipeline retry and cancel operations | Always | Guide 2 (PROTECTED) |
| Any command moving data across the containment boundary | Always | Guide 2 (PROTECTED) |
Every action on this list prompts you always, and the Profile column says which deployments carry it. The first four are the Guide 1 set that every deployment carries. The three shown in blue apply only where your team runs Guide 2, because at PROTECTED the containment boundary and the code platform both become things a human signs off rather than a classifier. Your team may have extended the set further, so treat the list as the floor rather than the whole of it.
Untrusted content
You handle untrusted content every day through merge request diffs, issue text, log files, dependency source and pipeline output. Four practices apply.
- 1Review suggested commands before you approve them. The prompt is the control, and skimming it defeats the control.
- 2Do not pipe untrusted content directly to Claude. Preprocess it first. Factor 09 covers the mechanism, which cuts cost as well as risk.
- 3Verify proposed changes to critical files. Migrations, infrastructure code and authentication logic warrant a read rather than a glance.
- 4Report suspicious behaviour with
/feedback.
The agent cannot fetch
The single highest-value deny rule on the baseline closes the most direct ingestion and exfiltration path for you. WebFetch, WebSearch, curl, wget and the PowerShell download commands are denied by managed rule. That removes the agent's fetch capability, not your endpoint's: your browser and your tooling are unaffected.
WebFetch against allowlisted domains for developer experience. The Bash download denies stay either way, because shell downloads are harder to constrain by destination. Check with your platform owner before assuming a fetch will work.
Detection sits behind the structural controls
Your organisation runs endpoint detection and response with AI agent runtime inspection, scanning the agent loop at your prompt, before each tool call, and on each tool response. In block mode a detected injection stops before the action runs, and you are notified both in the agent and by system notification. At PROTECTED, block mode is required rather than a maturity target, and the agent runs inside the contained environment where the inspection happens.
Boundaries, settings and credentials
managed-settings.json, which is read-only to you. It overrides anything you or a repository set, and managed deny rules cannot be overridden by a lower-tier allow rule. You cannot configure your way out of the posture, so where a rule blocks legitimate work, raise it through the request path rather than working around it.ConfigChange hook logs or blocks in-session settings changes, hooks are restricted to managed sources, and permission bypass is disallowed..env files, key files and credential directories. A Read deny on a sensitive path also prevents the editor selection and open-file notice for that file reaching the model, which serves content control and cost together.managed-mcp.json lists only the CCS GitLab MCP server, and the CLI loads nothing else: you cannot add servers, repository .mcp.json files are ignored, plugin servers are suppressed and --mcp-config is refused. An MCP server is remote code with tool authority over the agent, so it gets governed like software you install. A blocked add attempt shows an enterprise policy error, but a previously configured server that becomes blocked simply disappears without explanation. Propose a server through the request path rather than configuring around it.The editor context you are sending without noticing
Read deny rule for any path that should never reach the model.
This one server is the exception to the deny-by-default MCP position above. The extension's in-process server still loads in sessions the extension starts, regardless of managed-mcp.json, because it is product behaviour rather than a configurable setting. At PROTECTED the editor runs connected into the contained environment through remote-development tooling, which makes the boundary invisible in the editor but does not change this behaviour.
Security tooling available to you
Runs an on-demand security pass over the changes on your current branch. Worth running before any merge request touching authentication, input handling or data access.
Has Claude review and fix vulnerabilities in its own code changes during the session, rather than waiting for a review gate to find them.
Overall, the posture holds without your attention in most respects, and the two places it does not are the prompts you approve and the content you feed in.
Show answer
Prerequisites
Running claude in the integrated terminal requires the standalone CLI on your shell PATH. Installing the graphical extension does not provide it, because the extension bundles a private copy of the CLI for its own chat panel.
claude --version before your first session.
Without the PATH entry, six behaviours this guide teaches are unavailable, and each is a cost-reducing behaviour elsewhere in the training.
| Command | Purpose | Taught in |
|---|---|---|
claude | Launching the CLI at all, which is the supported configuration | Everywhere |
claude --permission-mode plan | Starting directly in plan mode | Factor 02 |
claude -p | Non-interactive runs, including pre-commit hook automation and fan-out loops over a file list | Factor 09 |
claude --resume, --continue | Returning to a named session instead of rebuilding context | Factor 05 |
claude mcp list | Confirming which MCP servers a session will load, the managed MCP validation check | Factor 10 |
claude --version | Version verification | Factor 12 |
claude -p, and do not plan a pipeline stage that calls the model.
/clear and /compact are in-session commands and do not depend on PATH, although starting the CLI does. If claude --version fails, raise it with your platform owner rather than installing the CLI yourself, since a user-installed copy will not satisfy application control.
The Three Work Patterns
Two foundational patterns cover secd3v Claude Code usage, with a third mode that modifies how one of them runs. Identify which one you are in before you apply any factor, because the same behaviour carries different value in each.
| Pattern | When it applies | Session shape | Dominant cost driver |
|---|---|---|---|
| DevSecOps | Existing production system with live users. Code review, bug fixes, security checks, merge request review under active security constraints | 8 to 11 short sessions per day, 5 to 45 min each. You approve every step | Cache reads across many short sessions, plus regular input from broad prompting |
| Agentic | Building something new. New service, new module, greenfield implementation where wrong output gets discarded cheaply | 1 to 2 long sessions per day, up to 4 hours. Claude works autonomously | Regular input at 25,000 tok per turn at heavy usage, 4.1 times the DevSecOps volume |
| Spec-Driven | Agentic variant. Author a specification before executing, then execute against it rather than discovering scope through file exploration | Phase 1 spec write 30 min, Phase 2 execution up to 4 hrs, Phase 3 conformance review 30 min | Regular input cut to approximately 13,000 tok per turn, a 48% reduction |
Why DevSecOps is the brownfield pattern
Brownfield work means an existing production system with live users, an established architecture, years of accumulated code, technical debt and security constraints under active enforcement. You cannot start fresh, cannot discard a wrong implementation without consequence, and cannot hold broad autonomous permissions without risk. In government, defence and high compliance contexts this describes the overwhelming majority of daily work.
DevSecOps is the direct response to those constraints rather than a methodology that happens to sit alongside them. Where a codebase has a security posture that cannot be accidentally degraded and a production environment where mistakes have immediate consequences, the human-gated, review-at-every-step workflow becomes mandatory rather than optional discipline.
That shapes how you use the CLI. You say "look at auth.py lines 42 to 89 and identify any SQL injection risk" rather than "explore this codebase and make the changes you think are needed". Sessions run short, targeted and bounded, because the work demands precision over autonomy.
Why agentic is primarily the greenfield pattern
Greenfield work builds a new service, module or application, with no live users to disrupt, no established architecture to break and no accumulated constraints to misunderstand. A wrong-direction implementation gets discarded cheaply, so the cost of an incorrect autonomous attempt stays low. Claude explores, plans and implements, reading many files and building its own context map over long sessions.
Agentic development also suits specific brownfield scenarios, including large-scale migration sprints, comprehensive test generation and automated documentation. Those cases require three conditions: plan mode before any execution, a git worktree for isolation, and every checkpoint treated as a safety gate.
Session shapes
Use these to recognise when a session has drifted out of its intended type. Factor 12 turns them into telemetry thresholds.
| Pattern | Type | Shape and typical work | Turns | Regular input / turn | Output / turn |
|---|---|---|---|---|---|
| DevSecOps | Micro | 5 turns, 5 to 10 min. Doc update, syntax check, single-function review, pipeline triage, quick explanation | 5 | 2,500 tok | 400 tok |
| DevSecOps | Standard | 9 turns, 15 to 25 min. Code review, targeted bug fix, test generation, single-file security check, MR feedback | 9 | 5,000 tok | 700 tok |
| DevSecOps | Extended | 15 turns, 30 to 45 min. Multi-file security audit, SAST triage, compliance check, refactoring plan and execute | 15 | 9,000 tok | 1,000 tok |
| Agentic | Light | 5 sessions of 10 turns. Small feature additions, single-module builds, focused bug fixes | 50 / day | 4,000 tok | 600 tok |
| Agentic | Medium | 2 sessions of 25 turns. Feature implementation, module construction, MR creation | 50 / day | 13,000 tok | 800 tok |
| Agentic | Heavy | 1 session of 65 turns plus 1 of 25. New service construction, large autonomous implementation from spec | 90 / day | 25,000 tok | 1,100 tok |
Spec-driven sprint days
A spec-driven sprint runs three day shapes rather than one. Recognising which day you are in tells you which model and effort level belong.
| Day | Session mix | Turns | What happens |
|---|---|---|---|
| Spec day | 1 spec-write, 1 execution start | ~25 | Phase 1 plus Phase 2 kickoff. Spec authored, execution begins |
| Exec day | 1 to 2 execution sessions | 65 to 90 | Phase 2 sustained. Building against the spec, /compact as needed |
| Review day | 1 conformance, 1 correction | ~20 | Phase 3 review. Deviation notes feed a Phase 2 correction turn |
Which factors pay in which pattern
- Factor 01 · Prompt specificity
- Factor 05 · Session hygiene
- Factor 06 · Model selection and effort
- Factor 09 · Preprocessing hooks
- Factor 02 · Plan mode
- Factor 03 · Verification targets
- Factor 04 · Spec-driven development
- Factor 08 · Subagents
Overall, the pattern determines the dominant cost driver: DevSecOps spends on cache reads across many short sessions, while agentic spends on regular input inside a small number of long ones.
Prompt Specificity & Context Front-Loading
The largest per-session cost lever you control. How you phrase a request determines how much context Claude reads before it can act, and every token read compounds across the rest of the session.
DevSecOps examples
Security review of an authentication function
Fixing a known bug
Merge request review with a security focus
Agentic example: a new API endpoint
The 5W1H checklist
Answer these six before sending. If you can answer them, write the answers into the prompt.
| Question | What it covers | Example in prompt |
|---|---|---|
| What | Exactly what to change, review or understand | "review the token validation logic" not "check auth" |
| Where | Specific file path and line numbers | "@src/auth/tokens.py lines 88-134" |
| Why | Error message, security concern, test failure | "returning 401 for valid tokens after migration" |
| Look for | Specific patterns or vulnerability class | "flag non-constant-time string comparisons" |
| Don't touch | Files or logic to leave unchanged | "do not modify the database schema or migrations" |
| How | Pattern or library to follow | "use the same approach as @src/auth/refresh.py line 44" |
Specificity works differently in each pattern
| Pattern | What specificity does |
|---|---|
| DevSecOps | Keeps sessions inside their session-type scope and stops a micro session drifting into extended territory |
| Agentic | Vague prompts trigger broad file scanning, and each read compounds into every subsequent turn of a long session |
| Spec-Driven | The spec is the specificity mechanism, but the execution prompt must still reference specific spec sections rather than leaving Claude to interpret the whole document |
Quick reference by task type
| Task | Instead of | Say |
|---|---|---|
| Security check | "check for security issues" | "review @api/views.py lines 55-90 for IDOR, ensure user_id is validated against request.user" |
| Test generation | "write tests for the payment module" | "write pytest unit tests for charge() in @payments/processor.py using fixtures from @tests/conftest.py, cover success, declined card, network timeout" |
| Documentation | "document the API" | "add Google-style docstrings to the 4 public methods in @src/api/client.py, type annotations already present, skip them" |
| Pipeline triage | "CI is failing, fix it" | "job 'test-unit' in pipeline #4821 fails with ImportError: cannot import 'TokenCache' from 'src.cache', class renamed last commit. Fix imports in test files only" |
| Refactoring | "refactor the database layer" | "extract retry logic from @db/connection.py lines 120-156 into a RetryPolicy class in @db/retry.py, keep the existing interface, callers must not change" |
When a vague prompt is the right call
What would you improve in this file? surfaces things you would not have thought to ask about. Use it deliberately, in a session you intend to clear, rather than as a default.
src/payments/processor.py but not the exact lines. Which prompt costs least?Show answer
Plan Mode Before Execution
Plan mode separates what Claude should work out from what it should build. Correcting direction at the plan stage costs roughly 500 tokens. Correcting after twenty turns of wrong implementation costs tens of thousands.
Shift+Tab before typing. There is no /plan command. Review the plan, edit it with Ctrl+G, then switch back to execute. Skip plan mode for small clearly scoped tasks where you could describe the diff in one sentence.How to enter and leave it
- 1Enter. Press
Shift+Tabuntil the status bar shows plan mode on, or start withclaude --permission-mode plan. The start flag depends on the CLI being on your PATH. - 2Explore. Claude reads relevant files and asks clarifying questions without touching your code.
- 3Review and edit. Press
Ctrl+Gto open the plan in your text editor. Correct the approach, add constraints, or reject it, all at near-zero cost. - 4Execute. Approve the plan or press
Shift+Tabto switch back. Claude implements against the plan it agreed.
When to use it
- The task touches multiple files or modules
- You are unfamiliar with the code being changed
- You are uncertain what the right approach is
- Any brownfield change to a production system
- New feature implementation in an existing service
- Refactoring with potential side effects
- Security-sensitive changes
- Spec-driven Phase 2, before each significant component
- You could describe the complete diff in one sentence
- Fixing a typo, renaming a variable, adding a log line
- Adding a docstring to a clearly understood function
- Updating a dependency version
- A formatting-only change
- Any task where the scope is fully clear upfront
DevSecOps example: multi-file bug fix
Agentic example: new service construction
@src/api/users.py. The logic checks that the email field matches a regex. Use plan mode?Show answer
Verification Targets
Claude stops when the work looks done. Without a check it can run, "looks done" is the only signal available, and you become the verification loop. That is the most expensive loop in the workflow, because every mistake waits for you to notice it.
What counts as a check
Anything returning a signal Claude can read in the conversation. A test suite, a build exit code, a linter, or a script that diffs output against a fixture.
| Instead of | Write |
|---|---|
"implement a function that validates email addresses" | "write a validateEmail function. Test cases: user@example.com is true, invalid is false, user@.com is false. Run the tests after implementing" |
"the build is failing" | "the build fails with this error: [paste error]. Fix it and verify the build succeeds. Address the root cause, do not suppress the error" |
"add retry logic to the client" | "add retry logic to @src/http/client.py. Write a test asserting three attempts on a 503 then a raise. Run the suite and show me the output" |
"make the dashboard look better" | "[paste screenshot] implement this design. Screenshot the result, compare it to the original, list the differences and fix them" |
Ask for evidence, not assertion
Three gate strengths, and what each costs
| Gate | Mechanism | Cost |
|---|---|---|
| In-prompt | Ask Claude to run the check and iterate in the same message | Nothing beyond the turns the iteration takes. Use this by default |
/goal condition | A separate evaluator re-checks the goal after every turn until it resolves | The evaluator runs every turn. Idle check-ins start a new turn carrying full context while background work runs. CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 is set unless goal-driven sessions are approved for your team |
| Stop hook | A script blocks the turn from ending until the check passes | Up to 8 consecutive blocks before Claude Code overrides the hook. Budget up to 8 extra turns per gated task, and keep the check cheap and fast |
/goal or a Stop hook only where a run needs to finish correctly without you watching it.
The second opinion
A verification subagent reviewing the diff in a fresh context is the cheapest quality gate available, because the agent doing the work is not the one grading it. Factor 08 covers the pattern and the prompt.
Show answer
Spec-Driven Development
The largest single return available in the cost model. Author a complete specification before execution begins, and cut heavy agentic spend by 28% for the price of one authoring session per sprint.
Why it matters beyond cost
Spec-driven development establishes the specification as the authoritative source of truth and treats code as a derivative artefact. It replaces loose, non-deterministic prompting with unambiguous, executable contracts grounded in architectural and security constraints. That supports the traceability the Australian Information Security Manual and ISO 27001 require, validating every change against documented requirements through automated gates in CI/CD. Security corrections propagate across future regeneration cycles, which mitigates architectural drift and preserves individual human accountability for AI-enabled outcomes.
The three-phase lifecycle
Why it works: the front-loading principle
The dominant agentic cost driver is regular input from file exploration and growing conversation history. In a standard agentic session Claude reads ten to fifteen files to discover scope before work begins, and each read compounds into every subsequent turn. A spec replaces that discovery phase with a single authored document. Claude reads the spec and the specific files it names, and nothing else.
The second saving is wrong-direction turns. A standard agentic session at heavy usage accumulates four to eight turns building something later discarded because a constraint was not understood upfront. A spec surfaces those errors at Phase 1 review, where correction costs roughly 500 tokens rather than tens of thousands.
Session sequence: clear between Phase 1 and Phase 2
Write the specification to a file before the authoring session ends, then start a fresh session to execute it. The execution session then carries clean context focused entirely on implementation and references the written spec. That sequence is also cheaper, because a fresh session carrying a 2,000-token spec file costs less than one carrying the full authoring history.
| Boundary | Rule |
|---|---|
| Phase 1 → Phase 2 | Write the spec to a file, end the session, start fresh referencing the spec path |
| Within Phase 2 | /compact at 80% context fill. Focus it: /compact focus on spec compliance and decisions made so far |
| Phase 2 → Phase 3 | Start fresh, referencing the spec file by path |
| Between sprints | /rename then /clear. A new sprint means a new Phase 1 |
Let Claude interview you to produce the spec
The most efficient route to a spec is to have Claude ask you the questions. Start Phase 1 with a minimal prompt and let the interview surface what you have not considered, which is precisely the wrong-direction cost the spec exists to remove.
What a good spec contains
A spec is not a project overview. It is a set of actionable constraints Claude can follow and verify against. Every line should be something Claude does differently because it is there. The most useful specs are self-contained: they name the files and interfaces involved, state what sits out of scope, and end with an end-to-end verification step that proves the feature works.
Where the spec lives
Current sprint spec: @SPEC.md.
CLAUDE.md above 3,000 tokens on a spec-driven project indicates this failure is active. Factor 10 covers the fix.
Why Opus 4.8 is justified at Phase 1 for every tier
Model access is normally gated by developer tier. Spec authoring creates an amortisation justification that applies regardless of tier. A single Opus 4.8 spec session producing a tight 1,500 to 2,500 token specification amortises across 40 to 65 Sonnet 5 execution turns, and the overhead recovers within the first execution session through reduced file-exploration turns and eliminated wrong-direction corrections.
Telemetry signals
- Phase 2 regular input below 15,000 tok per turn
- Cache hit rate above 80% in Phase 2
- CLAUDE.md below 3,000 tokens
- Opus usage confined to Phase 1
- Phase 3 on Haiku for 70% or more of turns
- Phase 2 regular input above 18,000 tok per turn, so Claude is still file-exploring
- CLAUDE.md above 3,000 tokens, so the spec is embedded
- Opus usage in Phase 2 or 3, which is policy drift
- More than 2 wrong-direction corrections per Phase 2 session, so revisit Phase 1
- Phase 3 running Sonnet for all turns
Show answer
Session Hygiene
Every API call processes the full conversation history to date. Accumulated irrelevant turns add tokens to every subsequent message and degrade output quality, because performance falls as the context window fills.
/clear between every unrelated task in DevSecOps work. After correcting Claude twice on the same issue, clear and start fresh with a better prompt. Prefer /clear over /compact, because compaction is itself a large request. Use /btw for questions that do not need to stay in context.The commands
| Command | What it does | When |
|---|---|---|
/clear | Resets the context window completely | Between unrelated tasks. After two failed corrections. Switching project or codebase |
/rename | Names the session so you can resume it | Before /clear, whenever you might return |
claude --resumeclaude --continue | Returns to a named session instead of rebuilding context | Work spanning multiple sittings. Treat named sessions like branches, one per workstream |
/compact | Summarises history, keeping code and decisions | Long agentic sessions at 70 to 80% context fill, where continuity matters |
/compact focus on X | Focuses the summary | Continuing with a specific subset of a long session |
/btw | Asks a side question whose answer never enters conversation history | Checking a detail without growing context. The cheapest question you can ask |
Esc Esc or /rewind | Rewind menu, with summarise-from-here and summarise-up-to-here | Condensing part of the conversation while leaving the rest intact |
/context | Shows what is currently consuming context | Before deciding whether to compact or clear |
/clear over /compact. /compact reads the conversation it summarises, so compacting a large context is itself a large request, and the following turn writes a fresh prefix because summarisation replaces the conversation prefix. /clear costs nothing. Reserve /compact for continuity within one task, and where you only need part of the conversation condensed, the rewind menu costs less than a full compaction.
DevSecOps: every task boundary is a clear boundary
/clear and restate the task with the correction built into the prompt. A clean session with a better prompt almost always outperforms a long session with accumulated corrections.
The 5-minute cache is the default, and it refreshes for free
Prompt caching is the largest cost lever in the service, cutting cost by 38 to 63% against an uncached baseline. It is enabled by default on the Bedrock API and the 5-minute TTL applies unless a request specifies otherwise. Two mechanics matter to how you work.
- 1The TTL resets on every hit, at no charge. The cache expires only if no hits occur within the window, and each use refreshes it for free. A Claude Code session is a tool-use loop in which one conversational turn issues many API requests seconds apart, so the gaps that matter are genuine idle periods rather than turn boundaries. Measured cache hit rates exceed 90% on the 5-minute default.
- 2Editing CLAUDE.md invalidates the cache. Where a change is a one-off constraint, state it in the prompt instead and update CLAUDE.md at the end of your working block.
ENABLE_PROMPT_CACHING_1H. It is a process-level environment variable, not a per-request setting, so it applies to every session you run and cannot be selected per session type. A 1-hour write costs 2.0 times base input against 1.25 times for a 5-minute write, and that 60% premium applies to every cache-write token rather than only the opening prefix. Enabling it costs a heavy DevSecOps developer $10 to $14 per month and a heavy agentic developer $3 to $4, for a benefit the telemetry does not show.
The narrow cases where a 1-hour profile is warranted
Five situations genuinely produce expiry. Each needs measured evidence before the service enables the profile for your developer account, because it is a per-developer-profile decision rather than a per-session one.
| Case | Why the 5-minute cache expires |
|---|---|
| Long single turns | The lifetime measures from the start of the request, not the end of the response, and generation time counts against it. A four-minute stream leaves about one minute for the follow-up. Opus 4.8 at high effort is the exposure, and it is the one interactive case that expires a 5-minute cache without you pausing at all |
| Human review gaps | Merge request review, compliance reading and multi-file approval, particularly in Manual permission mode |
| CI pipeline waits | GitLab MCP workflows that block on a pipeline |
| Cross-session cache grouping | Impossible on a 5-minute TTL. Bedrock uses organisation-level cache isolation, so sharing is available. Measure the benefit before claiming it |
| Rate limit headroom | Cache hits are not deducted against rate limits, so fewer re-writes buys throughput on a constrained quota |
The long context effect
Sonnet 5, Opus 4.8 and Sonnet 4.6 carry a 1M token context window at standard rates with no long-context surcharge. A larger window delays auto-compaction, so sessions accumulate more context before summarisation and per-turn regular input rises even though the per-token rate does not. Treat it as a token-volume effect and clear more often rather than relying on auto-compaction to arrive.
Persistent instructions belong in CLAUDE.md
If you restate the same instruction every session, such as "always use our custom logger, never print()", it belongs in CLAUDE.md rather than in your prompt. Instructions in CLAUDE.md survive /clear. Instructions in conversation history do not. The sprint spec, however, is not CLAUDE.md material: see Factor 04.
float() for currency amounts when your codebase requires Decimal(). What next?Show answer
Model Selection & Effort Levels
Defaulting to Sonnet for everything is the most common unnecessary cost. Haiku handles a quarter to a third of DevSecOps work and all spec-driven conformance review. Effort levels tune reasoning depth without changing model.
Escalate in the right order
Three levers answer a hard task, and model tier is the most expensive of them. The gateway enforces this ordering rather than leaving it to judgement.
- 1stEffort level. Sonnet 5 at high effort is the correct first escalation for a standard developer facing a hard task.
- 2ndSession profile. Plan mode, or a written spec, removes more wrong-direction cost than a larger model adds.
- 3rdModel tier. Opus 4.8 is the exception, not the default response to difficulty.
- 4thMax effort on Opus 4.8. Senior-approved sessions only, one turn at a time. The most expensive request available on this deployment.
Model selection guide
| Task type | Model | Why |
|---|---|---|
| Docstrings, comments, type annotations | Haiku 4.5 | Pattern completion, no deep reasoning needed |
| Pipeline failure triage and CI configuration | Haiku 4.5 | Log pattern matching, not novel analysis |
| Syntax fixes, formatting, boilerplate scaffolding | Haiku 4.5 | Mechanical transformation, no design decisions |
| Dependency version and compatibility lookups | Haiku 4.5 | Factual retrieval |
| GitLab issue summarisation | Haiku 4.5 | Text processing |
| Spec-driven Phase 3 conformance review | Haiku 4.5 | Pattern matching against defined criteria, not reasoning |
| Code review and security analysis | Sonnet 5 | Requires understanding of intent and edge cases |
| Bug fix analysis and implementation | Sonnet 5 | Reasoning about cause, effect and constraints |
| Test generation beyond trivial coverage | Sonnet 5 | Understanding behaviour under failure modes |
| Multi-file feature implementation | Sonnet 5 | Sustained multi-turn reasoning |
| Spec-driven Phase 2 execution | Sonnet 5 | Haiku eligible for simple bounded implementation turns |
| Known-pattern threat modelling, OWASP Top 10, CVE triage | Sonnet 5 | Established patterns rather than novel reasoning |
| Automated tasks with no human in the loop | Sonnet 5 | Nobody verifies the reasoning, so take the better cost-reliability trade |
| Spec-driven Phase 1 spec authoring | Opus 4.8, all tiers | Amortised across 40 to 65 Sonnet 5 execution turns |
| Novel exploit chain assessment | Opus 4.8, senior | Novel reasoning under genuine ambiguity |
| Compliance gap analysis against complex controls | Opus 4.8, senior | Multi-control reasoning, high downstream error cost |
| Greenfield architectural design | Opus 4.8, senior | One planning session prevents many wrong-direction turns |
What the models cost relative to each other
| Comparison | Rate ratio | Effective ratio for the same text |
|---|---|---|
| Haiku 4.5 against Sonnet 5 | 2.0 times | Approximately 2.6 times, since Haiku uses the previous tokenizer |
| Opus 4.8 against Sonnet 5 | 2.5 times | 2.5 times, since both use the new tokenizer |
| Opus 4.8 against Haiku 4.5 | 5.0 times | Approximately 6.5 times |
Why Sonnet 5 and Opus 4.8 rather than the 4.6 generation
| Comparison | Where the current model is better | Where it costs more |
|---|---|---|
| Sonnet 5 against Sonnet 4.6 | Anthropic positions it as its most agentic Sonnet. The rate card is 33% lower, and after the tokenizer increase measured spend is 13.3% lower for identical work, because input, output and both cache rates scale by the same factor | Nothing. Sonnet 5 costs less on the rate card and less in measured spend |
| Opus 4.8 against Opus 4.6 | Minimum cache checkpoint falls from 4,096 tokens to 1,024, which widens what can be cached in short sessions. Adaptive reasoning brings effort-level control. A new system instruction can be added partway through a conversation without invalidating the system or message caches | Approximately 30% more expensive in practice at an identical rate card, because of the tokenizer |
The Haiku caching constraint, in proportion
Switching model mid-session
Conversation history carries over and new API calls bill at the new model's rate. The service enforces model access by role-based routing, per user, per team and organisation wide, so a model outside your permitted tier is rejected. A rejection is policy working rather than a fault to raise.
Effort levels: four levels, set once per session
Effort controls how deeply Claude reasons, separately from which model you use. The resolved effort value renders into the prompt, so changing it between requests invalidates the message-block cache and forces a re-write. Choose the effort level with the session type at the start, then steer individual turns through prompt wording.
| Level | Command | Availability | When |
|---|---|---|---|
| Low | /effort low | All models | Routine work: docstrings, syntax fixes, formatting, quick lookups, conformance review |
| Medium (default) | /effort medium | All models | Most coding work: code review, bug fixes, standard implementation, agentic execution |
| High | /effort high | All models | Security analysis, architectural decisions, complex bug investigation, spec authoring |
| Max | /effort max | Opus 4.8 only | The most complex reasoning. Senior-approved sessions only, and it does not persist |
Match the level to the session type
| Session type | Effort |
|---|---|
| DevSecOps micro and standard | Low or medium |
| DevSecOps extended, agentic execution | Medium |
| Spec-driven Phase 3 conformance | Low |
| Spec-driven Phase 1 authoring, novel reasoning under ambiguity | High |
| Opus 4.8 on a single turn of genuinely novel reasoning, senior approved | Max |
Show answer
Extended Thinking
Thinking tokens bill as output tokens, the most expensive category. Anthropic enables extended thinking by default because it materially improves complex planning and reasoning. On the current generation it is also invisible, so you pay for reasoning you never see.
The control depends on the model generation
| Model | Control |
|---|---|
| Sonnet 5, Opus 4.8 | Adaptive reasoning only. Effort level is the control, and these models ignore a nonzero MAX_THINKING_TOKENS |
| Sonnet 4.6, Opus 4.6 | MAX_THINKING_TOKENS applies in combination with CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1. Relevant only on a validated legacy workload |
Since Sonnet 5 and Opus 4.8 are the models you will use for new work, effort level is your only lever. Factor 06 carries the four levels and the session-type table, and the rule to set it once and hold it applies here for the same cache reason.
/effort max runs on Opus 4.8 only, does not persist, and produces the largest thinking blocks available. Thinking bills as output at $27.50 per MTok on Opus 4.8, which is 12.5 times the Sonnet 5 input rate. Use it on one turn of a senior-approved session where the reasoning is genuinely novel, and never as a standing setting.
Thinking is invisible in the transcript
On Sonnet 5 and Opus 4.8, thinking display defaults to omitted rather than summarised. You will not see the reasoning and you still pay for it. Any cost estimate built from returned content under-reports on these models.
/cost at the end of a session and read the usage fields the Bedrock response reports. Factor 12 covers the telemetry.
Sonnet 5 at high effort against Opus 4.8
- The problem is complex but well defined
- You can specify the constraints clearly
- Cost matters
- This is your first escalation, always
- The reasoning is genuinely novel
- Security threat modelling with ambiguous attack surfaces
- Compliance analysis where multiple control interpretations exist
- Sonnet 5 has already failed twice at high effort
MAX_THINKING_TOKENS=2000 on a Sonnet 5 session to cap reasoning cost across a batch of routine tasks. What happens?Show answer
Subagents for Research and Cost Routing
Subagents run in separate context windows and report back summaries, so verbose exploration never enters your main conversation. They also serve as cost routers, letting a Sonnet 5 session buy Haiku-priced investigation.
model: haiku for file scanning, documentation lookup and log analysis. Subagents are enabled by default and need no configuration. A verification subagent reviewing the diff in a fresh context is the cheapest quality gate available.How to invoke one
What to route and what to keep
- Library and version compatibility checks
- Scanning for a pattern across many files
- Verifying test coverage for a module
- Checking whether a dependency is present
- Summarising a large file or log
- Running shell commands and returning results
- Work needing the full conversation context
- Implementation decisions depending on prior turns
- Security analysis requiring nuanced interpretation
- Anything where you need to review Claude's reasoning
- Anything modifying files in your project
The adversarial review subagent
A verification subagent reviewing the diff in a fresh context is the cheapest quality gate available, because it sees only the diff and the criteria rather than the reasoning that produced the change. The agent doing the work is not the one grading it.
Run the bundled /code-review skill for a correctness check on the current diff. To check the diff against your plan or spec instead, write the prompt yourself.
Because the reviewer runs as a subagent, the implementing session receives the gaps directly and can fix and re-review without you copying findings between windows.
Subagents against agent teams
| Single agent | Subagent | Agent team | |
|---|---|---|---|
| Context windows | 1 | 2, main plus subagent | 1 per member |
| Token multiplier | 1 times | ~1.2 to 1.5 times | ~7 times per member in plan mode |
| Enabled by default | Yes | Yes | No |
| Approval needed | No | No | Yes |
If you are unsure which you need, use a subagent. Cost Governance covers agent teams.
requirements.txt and at what versions. Cheapest approach?Show answer
Preprocessing Hooks
A hook preprocesses data before Claude sees it. Instead of Claude reading a 10,000-line log to find errors, a hook greps for what matters and returns only the matching lines, cutting context from tens of thousands of tokens to hundreds.
PreToolUse hook reduces the same read to hundreds of tokens, every time, without you remembering to ask.The mechanism
Hooks are guarantees, CLAUDE.md is a request
migrations/ is deterministic. Anywhere a control must happen every time with zero exceptions, use a hook rather than an instruction.
Claude will write them for you
Run /hooks to browse what is configured, and edit .claude/settings.json directly to configure by hand.
ConfigChange hook logs in-session settings changes. Raise a new hook through your team's configuration path rather than adding it locally, and expect it to be reviewed like any other managed control.
Where hooks pay most
| Situation | Hook |
|---|---|
| Verbose CI and pipeline output | Grep for error and failure patterns before the read reaches context |
| Large generated files that Claude occasionally reads | Truncate above a line threshold with a note explaining the truncation |
| A lint or type check that must run after every edit | PostToolUse hook returning only violations |
| A directory that must never be written | PreToolUse block, which is stronger than a CLAUDE.md rule |
| An unattended run that must not finish until tests pass | Stop hook, per Factor 03. Budget up to 8 extra turns |
Show answer
CLAUDE.md, Skills & Context Discipline
CLAUDE.md loads into every session automatically. Every token in it is billed at cache-read price on every API call in every session on that project. Keep it sharp: instructions Claude follows, not background Claude reads.
What belongs in CLAUDE.md
Run /init to generate a starter file from your project structure, then prune it. For each line, ask whether removing it would cause Claude to make a mistake. If not, cut it. Run /doctor on a checked-in file and Claude proposes cuts for content it can derive from the codebase.
- Bash commands Claude cannot guess
- Code style rules that differ from defaults
- Security rules, such as never logging PII
- Test patterns and required coverage
- Key file paths for common reference
- Forbidden patterns Claude should refuse
- Developer environment quirks
- Anything Claude can work out by reading the code
- Standard language conventions Claude already knows
- Detailed API documentation, link to it instead
- Why you chose a technology
- Architecture evolution narrative
- Team onboarding context
- Information that changes frequently
CLAUDE.md template, secd3v projects
Skills for anything only sometimes relevant
Domain knowledge and workflows that apply to some sessions belong in skills, not CLAUDE.md. Claude loads them on demand without bloating every conversation.
Skills also define repeatable workflows you invoke directly with /skill-name. Use disable-model-invocation: true for workflows with side effects you want to trigger manually.
The spec-driven bloat failure
/context, check the CLAUDE.md line, strip it to SPEC.md and leave a single reference line.
.claudeignore on day one
Configure .claudeignore on day one. Cover build outputs, dependency directories, generated files, binaries, large media and IDE directories.
MCP context and connection hygiene
The GitLab MCP server adds 9,000 tokens of tool names and schemas to the fixed cached context, and they load upfront on every session where the server is connected.
| Component | No MCP | With GitLab MCP |
|---|---|---|
| System prompt and agent instructions | 3,900 tok | 3,900 tok |
| Built-in tool definitions | 15,600 tok | 15,600 tok |
| CLAUDE.md and memory files | 1,500 tok | 1,500 tok |
| GitLab MCP tool names and schemas | not applicable | 9,000 tok |
| Total fixed cached context | 21,000 tok | 30,000 tok |
The invalidation risk here comes from the managed MCP configuration changing between sessions, not from loading a tool within one. Cache checkpoints chain in the order tools, then system, then messages, so a change to the tools section invalidates the system and messages caches with it.
| Scenario | MCP configuration |
|---|---|
| Pure coding session, no GitLab operations needed | Disconnect. You are paying 9,000 tokens of cached context per session for nothing |
| Merge request review, reading diffs, checking pipelines | Connect. The overhead earns its place |
| Mixed session with occasional GitLab checks | Connect, and consider whether a CLI covers the occasional check instead. See Factor 11 |
| A server you expected is missing | Check claude mcp list. managed-mcp.json lists only the CCS GitLab MCP server, so nothing else will connect |
GitLab calls also add per-turn cost beyond the fixed context
A GitLab call returns a payload that enters the conversation as an uncached tool result, whether that is a merge request diff, a pipeline log, an issue body or a review comment thread. That sits on top of the 9,000-token difference in fixed cached context.
| Pattern | Additional regular input per turn Est | Why |
|---|---|---|
| DevSecOps | ~1,010 to 1,122 tok | Merge request review and pipeline triage call GitLab on most turns |
| Agentic | ~450 to 700 tok | A long autonomous session touches GitLab less often relative to its turn count |
Output rises by up to 105 tokens per turn as well. These uplifts are the second-largest input to the with-MCP cost figures after the fixed context.
Show answer
CLI Tools & Code Intelligence Plugins
CLI tools are the most context-efficient way to reach an external service, because they add no per-tool listing to your context window. A code intelligence plugin replaces text search with precise symbol navigation.
--help. For typed languages, install a code intelligence plugin so one go-to-definition call replaces a grep and several speculative file reads.Teach Claude a tool it does not know
Claude is effective at learning CLI tools from their own help output, so an unfamiliar internal tool is not a reason to reach for an MCP server.
Code intelligence for typed languages
Run /plugin to browse the marketplace. Plugins bundle skills, hooks, subagents and MCP servers into a single installable unit.
managed-mcp.json position in Factor 10 exists for the same reason: an MCP server is remote code with tool authority over the agent, and it gets governed like software you install.
Show answer
Token Telemetry
Without instrumentation, optimisation is guesswork. Run /context at the start of any session where cost looks wrong, and /cost at the end of any session to see where the money went.
Health metrics
| Signal | Indicates | What to do |
|---|---|---|
| Cache hit rate below 70% | Poor session hygiene, or one of the invalidation causes below | Diagnose before changing any TTL setting. Below 50%, act immediately |
| Regular input above 15,000 tok per turn, DevSecOps | Broad prompting, or a missing /clear | Factor 01 and Factor 05 |
| Regular input above 18,000 tok per turn, spec-driven Phase 2 | The spec is not being referenced and Claude is still file-exploring | Tighten the spec scope, per Factor 04 |
| CLAUDE.md above 3,000 tokens on a spec-driven project | Spec prose embedded in CLAUDE.md | Extract to SPEC.md. This is an endpoint-side check only, since the gateway records metadata and cannot inspect the file |
| Haiku share below 15% on DevSecOps | Model discipline not applied | Factor 06 |
| Cache write records reporting a 1-hour TTL | The 1-hour setting is active on a profile that may not warrant it | Raise it with your platform owner |
| Cache token counts at zero | A prompt below the minimum cacheable length for the model in use | Almost always a very small subagent call, per Factor 06 |
Diagnose a cache miss before you change anything
A longer TTL fixes expiry. It does nothing for invalidation and costs the premium anyway. Attribute the miss first.
| Evidence | Cause |
|---|---|
| Full-prefix cache write after an idle gap | TTL expiry. The only evidence that justifies enabling the 1-hour profile for your account |
| Full-prefix cache write with no preceding gap | Invalidation. Work through the six causes below. The remedy is configuration discipline rather than a longer TTL |
| Cache write volume rising with no change in session pattern | Routing behaviour under load where a multi-region profile is in use |
The six invalidation causes
| Cause | Effect | Remedy |
|---|---|---|
| Effort level changed mid-session | The resolved effort value renders into the prompt, so a change invalidates message blocks | Set effort once per session and hold it, per Factor 06 |
| Thinking configuration changed | Same mechanism as effort | Set once per session |
| Tools section changed | Checkpoints chain in the order tools, system, messages, so modifying tools invalidates the system and messages caches with it | Hold the managed MCP configuration stable between sessions |
| Image added | Adding an image anywhere in the prompt invalidates message blocks | Expect a full re-write after a pasted screenshot |
| 20-block lookback exceeded | Automatic prefix checking looks back approximately 20 content blocks from the checkpoint, and static content beyond that range is not found | Occurs on long parallel tool sequences. Additional checkpoints are the documented remedy, up to the four-checkpoint maximum |
| Cross-region routing under load | At times of high demand, cross-region inference optimisations may lead to increased cache writes | Raise a sustained rise with your platform owner. Direct in-region routing is preferable for cache-sensitive workloads where capacity allows |
/compact also produces a full cache write on the following turn, because summarisation replaces the conversation prefix. That is a cost of compaction rather than a cache fault.
Reading a /context breakdown
The fixed cached context is approximately 21,000 tokens without MCP and 30,000 with the GitLab MCP server connected. Factor 10 carries the component breakdown.
| Signal in /context | Meaning |
|---|---|
| CLAUDE.md well above 1,500 tokens | Bloated. Review and trim per Factor 10 |
| CLAUDE.md above 3,000 tokens on a spec-driven project | Spec embedded. Extract to SPEC.md |
| Regular input far above the session-type figure | Session drift. Consider /clear |
| GitLab MCP schemas present when you have no GitLab work | 9,000 tokens of cached context you are paying for and not using. Disconnect |
| SPEC.md absent during Phase 2 | Phase 2 running without the spec, so the saving is not being realised |
Thinking is invisible, so do not estimate from the transcript
On Sonnet 5 and Opus 4.8, thinking display defaults to omitted. Any cost estimate built from returned content under-reports. Use /cost and read the cache and token counts the Bedrock response reports.
Where the numbers come from
Claude Code calls Bedrock directly and sends no usage metrics to Anthropic, so attribution comes from the request path. Every inference request passes through the CCS gateway, which is therefore the complete point of measurement: it prices and attributes each request, aggregated per user, per team, per model and organisation wide. The gateway also assumes its upstream role per user with session tags, so the AWS audit trail corroborates the gateway record rather than substituting for it.
/context at turn 5 of a DevSecOps session: CLAUDE.md 8,200 tokens, regular input 38,000 tokens, cache hit rate 45%. What are the two urgent problems?Show answer
Cost Governance
Two capabilities multiply cost rather than reduce it, and both need approval before use. Six further behaviours consume tokens outside the session shapes this guide describes, and three of them bill a full-context turn while you are doing nothing.
Agent teams
| Single agent | Subagent | Agent team | |
|---|---|---|---|
| Context windows | 1 | 2, main plus subagent | 1 per member |
| Token multiplier | 1 times | ~1.2 to 1.5 times | ~7 times per member in plan mode |
| Enabled by default | Yes | Yes | No, gated behind CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 |
| Approval needed | No | No | Yes, granted to your engineering team, with a service budget ceiling |
| Suits | All work | Research, verification, cost routing | Genuinely parallel independent tasks only |
Running an agent team: four controls
These apply once your engineering team has been granted approval and you are about to spawn an agent team. Every "teammate" below is a Claude instance, not a colleague.
- 1Run agent teammates on Sonnet 5. Never Opus 4.8. The multiplier applies to whatever rate you set, so the model choice compounds across every member.
- 2Keep the agent team small. Each agent teammate is a full multiplier rather than a marginal addition, so a fourth member costs what the first one did.
- 3Write focused spawn prompts. Agent teammates load CLAUDE.md, MCP servers and skills automatically, so a broad spawn prompt pays for that context on every member.
- 4Shut agent teammates down when their work is done. An idle agent teammate keeps consuming tokens. Leaving a three-member team running overnight is a significant unintended cost event.
- Three independent microservices simultaneously
- Test suites for separate unrelated modules
- Documentation across disconnected packages
- Migration scripts for independent database tables
- Any work with no shared state
- Sequential tasks disguised as parallel ones
- Tasks sharing files or modules
- Work where output of A feeds input of B
- Feature implementation, use plan mode and Sonnet 5
- Any task a single focused session could handle
/batch
/batch splits a change across 5 to 30 subagents, each in its own worktree, each opening a merge request. Treat it as an uncapped multiplier alongside agent teams, governed at the service budget rather than by guidance. Test the instruction on two or three files before running it at scale.
Background token consumers
Six behaviours consume tokens outside the session shapes in The Three Work Patterns, and three of them bill a full-context turn while you are doing nothing.
| Driver | Mechanism | Control |
|---|---|---|
| Compaction | /compact reads the conversation it summarises, so compacting a large context is itself a large request, and the following turn writes a fresh prefix | Prefer /clear between unrelated tasks. Reserve /compact for continuity within one task |
/goal conditions | A separate evaluator re-checks the goal after every turn, and idle check-ins start a new turn carrying full context while background work runs | CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 where goals are not required |
| Stop hooks | Blocks the turn from ending until the check passes, for up to 8 consecutive blocks | Keep the check cheap and fast. Budget up to 8 extra turns per gated task |
| Scheduled tasks | Fires on its interval and sends full context whether or not the session is active | Disabled by policy unless a named use case justifies it |
| Cross-session messaging | A message from another session arrives as a new turn with full context | crossSessionInbound set to hold |
Agent teams and /batch | Multipliers rather than overheads | Approval and a service budget ceiling |
Settings you should not change
Four service settings carry the estate's financial risk. They are set for you, and knowing why helps you recognise a misconfiguration on your own endpoint.
| Setting | Requirement | Consequence if omitted |
|---|---|---|
ENABLE_PROMPT_CACHING_1H | Unset, so the Bedrock 5-minute default applies. Enabled per developer profile only on measured expiry | Enabling it estate-wide costs $3 to $14 per developer per month for no measured benefit |
CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 | Set unless goal-driven sessions are an approved workflow | Idle check-ins send full context on an interval |
crossSessionInbound = hold | Set | Uncontrolled turns on idle sessions |
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS | Unset unless your engineering team holds approval, with a service budget ceiling | Approximately 7 times token usage |
managed-mcp.json | Lists only the CCS GitLab MCP server, and held stable between sessions | Uncontrolled MCP context and tool authority, plus cache invalidation when the tools section changes |
defaultMode | Pinned to auto mode through managed settings, with the organisational risky action set held as ask rules | Sessions fall back to Manual mode, which raises turn counts and monthly cost for identical work and reintroduces approval fatigue on trivia |
Overall, the agent teams flag carries the largest single financial risk at roughly 7 times token usage, and the caching default carries the largest recurring one.
Model & Planning Reference
A lookup rather than a read. Which models you have, how each one caches, why token counts differ between generations, and the planning allowance for work these figures exclude. Rates, monthly cost tables and the model access splits live in the companion cost model.
The models provided
Five models are available through the service, and the gateway enforces which of them your role can reach. Sonnet 5 and Opus 4.8 are the current generation and the right choice for any new work. Sonnet 4.6 and Opus 4.6 exist for workloads already built and validated against them.
| Model | Role | Min tokens per checkpoint | Max checkpoints | Supported TTL | Tokenizer |
|---|---|---|---|---|---|
| Haiku 4.5 | Routine tasks, subagents, spec-driven conformance review | 4,096 | 4 | 5 minutes, 1 hour | Previous generation |
| Sonnet 5 | Primary workhorse. Code review, analysis, bug fixes, spec execution | 1,024 | 4 | 5 minutes, 1 hour | Current |
| Sonnet 4.6 | Validated legacy workloads only | 1,024 | 4 | 5 minutes, 1 hour | Previous generation |
| Opus 4.8 | Maximum intelligence. Senior-approved, and spec authoring for all tiers | 1,024 | 4 | 5 minutes, 1 hour | Current |
| Opus 4.6 | Validated legacy workloads only | 4,096 | 4 | 5 minutes, 1 hour | Previous generation |
How the cache limits behave
Three mechanics govern the limits in the table above, and each has a consequence you can see in your own telemetry.
Why token counts differ between generations
The current generation uses a newer tokenizer that produces approximately 30% more tokens for the same text. In this catalogue that separates Sonnet 5 and Opus 4.8 from Sonnet 4.6, Opus 4.6 and Haiku 4.5. Rate cards do not reflect it, so the effect lands entirely on token volume and stays invisible in a price comparison.
| Comparison | Token difference | Net effect on spend |
|---|---|---|
| Sonnet 5 against Sonnet 4.6 | approx. +30% | Sonnet 5 costs approximately 13% less for identical work, because its lower rate card more than offsets the extra tokens |
| Opus 4.8 against Opus 4.6 | approx. +30% | Opus 4.8 costs approximately 30% more in practice, because the rate card is identical and only the token count moves |
| Haiku 4.5 | unchanged | Unchanged. Haiku keeps the previous-generation tokenizer, which is why it is cheaper against Sonnet 5 than the rate cards alone suggest |
Work these figures exclude, and the 10% allowance
Every cost figure for this service models software development: reading and writing code, reviewing changes, generating tests, auditing for security defects, and building services against a specification. The session shapes in The Three Work Patterns are drawn from that work and nothing else.
You will also use the agent for content activities that sit alongside development and are not costed anywhere: long-form documentation, release notes and changelogs, merge request and release summarisation, architecture and sequence diagrams, commit message drafting, incident write-ups, runbooks, onboarding material, and open-ended explanations of unfamiliar code.
One such session per day, modelled at 6 turns with 4,000 tokens of regular input and 1,500 tokens of output per turn on the standard developer split, adds the following proportion to the base figure for each pattern and tier.
| Pattern and tier, standard policy | One session / day | Two sessions / day |
|---|---|---|
| DevSecOps light, no MCP | 16.0% | 32.1% |
| DevSecOps heavy, no MCP | 8.0% | 15.9% |
| DevSecOps heavy, GitLab MCP | 6.3% | 12.5% |
| Agentic heavy, no MCP | 3.4% | 6.9% |
| Agentic heavy, GitLab MCP | 3.2% | 6.4% |
| Use | When |
|---|---|
| 10% | The default. One to two complementary sessions per developer per day, mid-range across all patterns and tiers |
| 15% | Where merge request summarisation, release notes or documentation form a stated part of the role. Also for developers at the light DevSecOps tier, where the allowance is proportionally larger against a smaller base |
| 5% | Defensible for a purely agentic team, because a heavy autonomous session dwarfs a short content session |
The allowance is a planning figure rather than a measured one. It gets replaced at the 30-day telemetry review by separating sessions whose output-to-input ratio exceeds roughly 1 to 3, which is the signature of content work rather than code work. If your own usage is mostly documentation and summarisation, say so when your team sets budgets, because the base figures will understate you.
Factor quick summary
| # | Factor | Impact | Key action |
|---|---|---|---|
| 01 | Prompt specificity | Highest | File path, line range, specific concern. 5W1H checklist. One task per message |
| 02 | Plan mode | Highest agentic | Shift+Tab or claude --permission-mode plan. Skip for one-sentence diffs |
| 03 | Verification targets | High | Give Claude a check that returns pass or fail. Ask for evidence, not assertion |
| 04 | Spec-driven development | Highest agentic | Author the spec, write it to SPEC.md, end the session, execute fresh. Phase 3 on Haiku |
| 05 | Session hygiene | High DevSecOps | /clear between unrelated tasks. Prefer /clear over /compact. /btw for side questions |
| 06 | Model selection and effort | High DevSecOps | Escalate effort, then session profile, then model. Set effort once per session |
| 07 | Extended thinking | High if unmanaged | Effort level is the control on Sonnet 5 and Opus 4.8. Thinking is invisible and still billed |
| 08 | Subagents | High agentic | Haiku subagents for research. Adversarial review subagent on the diff |
| 09 | Preprocessing hooks | High DevSecOps | Cut a 10,000-line log to hundreds of tokens before Claude reads it |
| 10 | CLAUDE.md, skills and context | Medium | Under 200 lines. Skills for sometimes-relevant knowledge. Spec to SPEC.md, never CLAUDE.md |
| 11 | CLI tools and plugins | Medium | Prefer a CLI over an MCP server. Code intelligence plugin for typed languages |
| 12 | Token telemetry | Foundation | /context to diagnose, /cost to attribute. Diagnose invalidation before blaming expiry |