Enterprise cost model · v3.0 · 23 August 2026
secd3v Claude Code Service: Cost Considerations

Enterprise cost considerations for DevSecOps (Brownfield), Agentic (Greenfield) and Spec-Driven development using the Claude Code CLI with the secd3v Claude Code Service, with or without GitLab via the secd3v GitLab MCP service. Covers the AU rate card and cache mechanics, the model access policy, prompt caching and TTL selection, developer optimisation, mandatory service configuration, and how spec-driven workflows modify token consumption, session shape and model allocation within the Agentic pattern. Every inference request reaches Amazon Bedrock in ap-southeast-2 (Sydney) and ap-southeast-4 (Melbourne) through the CCS gateway.

ModelsHaiku 4.5 · Sonnet 5 · Sonnet 4.6 · Opus 4.6 · Opus 4.8
PathAgent · CCS · AWS Bedrock · Claude Model
EndpointAU cross-region / in-region, +10%
Cache5-min TTL default
CurrencyUSD
Section 01

Executive Summary

$66DevSecOps heavy dev
std policy / month
$151Agentic heavy dev
std policy / month
$109Agentic spec-driven heavy
estimated / month
38–63%Cost reduction from
prompt caching

secd3v Claude Code Service usage in high compliance environments organises into two foundational patterns, DevSecOps (Brownfield) and Agentic (Greenfield), with a third mode, Spec-Driven Development, that modifies how Agentic sessions run. All three operate across Haiku 4.5, Sonnet 5 and Opus 4.8, with Sonnet 4.6 and Opus 4.6 retained for validated legacy workloads. Every cost carries the 10% premium that Australian endpoints attract, which is the price of keeping inference onshore and is not optional for a customer operating to the ISM PROTECTED baseline.

Prompt caching is the largest single lever, cutting cost by 38 to 63% against an uncached baseline. TTL selection within it is a second-order decision that favours the 5-minute default, and section 07 sets out the evidence. Spec-driven development adds a structural second lever. It front-loads reasoning into a stable cached specification, cuts regular input per execution turn by approximately 48%, and removes the most expensive tail-cost scenario in the model, the wrong-direction agentic session.

These figures cover software development work only. Section 06 sets out the complementary content activities they exclude and the allowance to add for them.

Overall, a heavy agentic developer costs $151 per month and a heavy DevSecOps developer $66 before that allowance, provided the service enforces the configuration in section 10.

Section 02

Platform & Pricing

secd3v pins all inference to Australian Amazon Bedrock endpoints in ap-southeast-2 (Sydney) and ap-southeast-4 (Melbourne), reached through the AU cross-region inference profile or, in Melbourne, through direct in-region routing. Both options carry a 10% premium over the global endpoint, and that premium is included in every figure in this document. Anthropic's Bedrock documentation states the premium and names the AU inference profile as the mechanism for routing across Australian regions within the geography.

Claude Haiku 4.5
Routine tasks · Subagents · Conformance review
Input$1.10 / MTok
Output$5.50 / MTok
Cache write 5-min$1.375 / MTok
Cache write 1-hr$2.20 / MTok
Cache read$0.11 / MTok
Min per checkpoint4,096 tok
Previous-generation tokenizer
Claude Sonnet 5
Primary workhorse · Code review · Execution
Input$2.20 / MTok
Output$11.00 / MTok
Cache write 5-min$2.75 / MTok
Cache write 1-hr$4.40 / MTok
Cache read$0.22 / MTok
Min per checkpoint1,024 tok
New tokenizer, approx. +30% tokens
Claude Opus 4.8
Maximum intelligence · Senior-approved · Spec-write exception
Input$5.50 / MTok
Output$27.50 / MTok
Cache write 5-min$6.875 / MTok
Cache write 1-hr$11.00 / MTok
Cache read$0.55 / MTok
Min per checkpoint1,024 tok
New tokenizer, approx. +30% tokens

AU rate card

Per million tokens (MTok), on-demand rates, Australian endpoints. Rates derive from Anthropic's published rate card with the 10% Australian endpoint premium applied. Confirm the effective rate in the AWS console, or through the AWS Price List API, before using these figures for contracted forecasting.

↔ scroll if needed
ModelRoleInputOutput5-min write1-hr writeCache read
Haiku 4.5Routine tasks, subagents, conformance review$1.10$5.50$1.375$2.20$0.11
Sonnet 5Primary workhorse$2.20$11.00$2.75$4.40$0.22
Sonnet 4.6Validated legacy workloads only$3.30$16.50$4.125$6.60$0.33
Opus 4.8Maximum intelligence, senior-approved$5.50$27.50$6.875$11.00$0.55
Opus 4.6Validated legacy workloads only$5.50$27.50$6.875$11.00$0.55
Cache pricing follows fixed multipliers on base input: a 5-minute write costs 1.25 times, a 1-hour write costs 2.0 times, and a cache read costs 0.1 times. The multipliers stack with the Australian endpoint premium rather than replacing it.

Per-model cache limits

AWS publishes the operative limits for the Bedrock API, and all five models in the catalogue support both TTL options with up to four cache checkpoints per request.

ModelMin tokens per checkpointMax checkpointsSupported TTL
Haiku 4.54,09645 minutes, 1 hour
Sonnet 51,02445 minutes, 1 hour
Sonnet 4.61,02445 minutes, 1 hour
Opus 4.81,02445 minutes, 1 hour
Opus 4.64,09645 minutes, 1 hour
Three mechanics govern how those limits behave, and each has a cost consequence.
  • The minimum is cumulative, not per section – AWS evaluates the minimum against the combined token count across the tools, system and messages sections rather than each section individually. With 21,000 tokens of fixed context in every session, the 4,096-token minimum on Haiku 4.5 and Opus 4.6 is comfortably exceeded, so the threshold constrains only a subagent invocation carrying a genuinely small prompt and no tool definitions.
  • Sections chain, so an early change invalidates everything after it – checkpoints process in the order tools, then system, then messages. Modifying the tools section invalidates the system and messages caches with it.
  • Inference will succeed silently when caching does not – a request below the minimum still returns a response, with no error, and nothing cached. Cache token counts at zero are the only signal.

The tokenizer difference between generations

Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, and Sonnet 4.6 and earlier use the previous one. In the current catalogue this separates Sonnet 5 and Opus 4.8 from Sonnet 4.6, Opus 4.6 and Haiku 4.5. Rate cards do not reflect it, so the effect lands entirely on token volume and stays invisible in a price comparison.

ComparisonRate differenceToken differenceNet effect on spend
Sonnet 5 against Sonnet 4.6$2.20/$11.00 against $3.30/$16.50approx. +30%Sonnet 5 approx. 13% cheaper
Opus 4.8 against Opus 4.6Identical at $5.50/$27.50approx. +30%Opus 4.8 approx. 30% more expensive
Sections 04 and 06 apply the 1.30 multiplier to Sonnet 5 and Opus 4.8. Verify the multiplier against production telemetry within the first 30 days and restate it, because the exact increase depends on content and workload shape.

Endpoints and regions

Both Australian regions serve the catalogue, and they differ in how a request reaches them. Sydney routes through the AU cross-region inference profile. Melbourne supports the AU profile and also direct single-region routing without a profile. Both attract the 10% premium.

Cross-region inference can raise cache-write volume on its own. AWS states that cross-region inference selects the optimal region within the geography to serve a request, and that at times of high demand those optimisations may lead to increased cache writes. A cache write costs 1.25 times base input against 0.1 times for a read, so a period of high regional demand raises cost with no change in developer behaviour.

CCS supports Australian cross-region and in-region routing only, so a request may move between ap-southeast-2 and ap-southeast-4 under the AU profile and will never leave Australia. The cache-write behaviour above therefore applies within Australia and will impact costs. Treat an unexplained cache-write rise as a routing symptom.

Context window and quotas

Sonnet 5, Opus 4.8 and Sonnet 4.6 carry a 1M token context window at standard rates with no long-context surcharge. A larger window delays auto-compaction, so sessions accumulate more context before summarisation and per-turn regular input rises even though the per-token rate does not. Treat this as a token-volume effect rather than a pricing effect.

The default Bedrock quota is 2 million input tokens per minute, extensible to 4 million on request without additional Anthropic approval, with requests-per-minute limits enforced separately by AWS. Cache hits are not deducted against rate limits, which makes caching a throughput lever as well as a cost lever.

Overall, the Australian rate card carries a 10% sovereignty premium and a 30% tokenizer effect on the two current-generation models, and cross-region routing can raise cache-write volume independently of developer behaviour.

Section 03

Development Patterns & Session Definitions

Why DevSecOps is a brownfield pattern

Brownfield development means working on an existing production system with live users, an established architecture, years of accumulated code, technical debt and security constraints under active enforcement. A developer cannot start fresh, cannot discard a wrong implementation without consequence, and cannot hold broad autonomous permissions without risk. In government, defence and high compliance contexts this describes the overwhelming majority of daily work.

DevSecOps is the direct response to those constraints rather than a methodology that happens to sit alongside them. Where a codebase has a security posture that cannot be accidentally degraded and a production environment where mistakes have immediate consequences, the human-gated, incremental, review-at-every-step workflow becomes mandatory rather than optional discipline.

This shapes how the CLI gets used. The developer says "look at auth.py lines 42 to 89 and identify any SQL injection risk" rather than "explore this codebase and make the changes you think are needed". Sessions run short, targeted and bounded, because the work demands precision over autonomy.

Why Agentic is primarily the greenfield pattern

Greenfield development builds a new service, module or application, with no live users to disrupt, no established architecture to break and no accumulated constraints to misunderstand. A wrong-direction implementation gets discarded cheaply, so the cost of an incorrect autonomous attempt stays low. Claude explores, plans and implements, reading many files and building its own context map over long sessions.

Agentic development also suits specific brownfield scenarios, including large-scale migration sprints, comprehensive test generation and automated documentation, provided the developer enters plan mode before any execution, works in a git worktree for isolation, and treats every checkpoint as a safety gate. It does not suit routine brownfield maintenance, security operations, or any work where an incorrect autonomous change reaches production. The session risk profile determines the fit, not the label.

Spec-Driven Development as a structured variant SDD

Spec-driven development overlays the Agentic pattern. The developer authors a structured specification covering interface contracts, data shapes, acceptance criteria, file layout, security constraints and test coverage, then executes against the spec rather than exploring freely.

SDD front-loads the reasoning cost. Reasoning concentrates in a short spec-writing session and every subsequent execution session becomes cheaper and more predictable. Wrong-direction errors surface at spec review, costing roughly 500 tokens to correct, rather than after twenty agentic turns. Regular input per execution turn falls by approximately 48%.

Spec-driven development matters in high-compliance, government and defence work because it establishes the specification as the authoritative source of truth and treats code as a derivative artefact. It replaces loose, non-deterministic prompting with unambiguous, executable contracts grounded in architectural and security constraints. It supports the traceability the Australian Information Security Manual (ISM) and ISO 27001 require, validating every change against documented requirements through automated gates in CI/CD. Security corrections propagate across future regeneration cycles, which mitigates architectural drift and preserves individual human accountability for AI-enabled outcomes in mission-critical environments.

Phase 1 · Spec Writing
Author before executing
Interface contracts, acceptance criteria, file layout, security constraints, test scope. Quality here has compounding leverage.
Opus 4.8, justified for all tiers
8–15 turns · 15–30 min
Regular input approximately 4,000 tok/turn
Output: 1,500–2,500 token spec
→
Phase 2 · Spec Execution
Implement against spec
The spec replaces file-discovery turns. Regular input approximately 48% lower than standard agentic. Start a fresh session; the spec lives in a file.
Sonnet 5 primary Haiku 4.5 for simple impl turns
40–65 turns · up to 4 hrs
Regular input approximately 13,000 tok/turn
/compact at 80% context fill
→
Phase 3 · Conformance Review
Check output vs spec
Pattern matching against spec criteria. Failures return to Phase 2 with specific deviation notes.
Haiku 4.5 primary Sonnet 5 for edge cases
5–10 turns · 15–30 min
Regular input approximately 3,000 tok/turn
Starts fresh, references the spec file
Session sequence. Write the specification to a file before the authoring session ends, then start a fresh session to execute it. The execution session then carries clean context focused entirely on implementation and references the written spec. That sequence is also cheaper, because a fresh session carrying a 2,000-token spec file costs less than one carrying the full authoring history.

Pattern contrast

DevSecOps (Brownfield)
  • Existing production system – live users, established architecture, active security constraints, technical debt
  • 8–11 sessions/day, 5–45 min each – short, targeted, bounded by task scope
  • Developer provides explicit file refs and line ranges; Claude does not explore freely
  • Human approves every response before Claude proceeds; accountability is non-negotiable
  • 5–15 API turns per session · avg 2,500–9,000 tokens regular input/turn
  • Surgical output: patches, analysis, test stubs at 400–1,000 tokens/turn
  • Plan mode before any multi-file change, mandatory practice
  • Cache writes amortise poorly per session, so cross-session behaviour matters
  • Haiku suitable for 25–30% of interactions
Agentic (Greenfield), standard and spec-driven
  • New codebase under construction – no live users, wrong implementations cheap to discard
  • Standard: 1–2 long sessions/day · Spec-driven: Ph1 (30 min) + Ph2 (up to 4 hrs) + Ph3 (30 min)
  • Standard: Claude explores freely at 25,000 tok avg regular input/turn
  • Spec-driven: Claude executes against the spec at approximately 13,000 tok avg regular input/turn (−48%)
  • 50–65 API turns/heavy session · auto-compaction at approximately 80% context fill
  • Spec-driven: single Opus spec session amortised across 40–65 Sonnet 5 execution turns
  • Cache writes amortise well within the session; the spec improves the hit rate further
  • Standard: Haiku 10–15% · Spec-driven: Haiku rises to ~20%

Session type definitions: DevSecOps Brownfield

Micro 5 turns
5 turns · 5–10 min
Doc update, syntax check, single-function review, pipeline triage, quick explanation
API turns5
Avg regular input/turn2,500 tok
Avg output/turn400 tok
Effort levellow / medium
Standard 9 turns
9 turns · 15–25 min
Code review, targeted bug fix, test generation, single-file security check, MR feedback
API turns9
Avg regular input/turn5,000 tok
Avg output/turn700 tok
Effort levellow / medium
Extended 15 turns
15 turns · 30–45 min
Multi-file security audit, SAST triage, compliance check, refactoring plan + execute
API turns15
Avg regular input/turn9,000 tok
Avg output/turn1,000 tok
Effort levelmedium

Session type definitions: Agentic Greenfield

Light 5 sessions of 10 turns
5 sessions · 10–20 min each
Small feature additions, single-module builds, code explanations, focused bug fixes
API turns/day50
Avg regular input/turn4,000 tok
Avg output/turn600 tok
Medium 2 sessions of 25 turns
2 sessions · 30–60 min each
Feature implementation, module construction, multi-file build, MR creation
API turns/day50
Avg regular input/turn13,000 tok
Avg output/turn800 tok
Heavy 65 turns + 25 turns
1 heavy session, 1 medium session
New service construction, greenfield architecture, large autonomous implementation from spec
API turns/day90
Avg regular input/turn25,000 tok
Avg output/turn1,100 tok

Daily usage tier definitions

↔ scroll if needed
TierSession mix / daySessionsTurnsDeveloper profile
DevSecOps
Light5 micro, 3 standard852Part-time AI assistance; quick reviews and consultations
Medium3 micro, 4 standard, 1 extended866Active developer; daily code review, security checks, bug fixes
Heavy5 micro, 4 standard, 2 extended1191Lead developer / security engineer; audits, MR reviews, compliance
Agentic, standard
Light5 light sessions550Light autonomous tasks; feature additions, focused builds
Medium2 medium sessions250Active builder; feature implementation, module construction
Heavy1 heavy session, 1 medium session290Power developer; new service construction, long autonomous sessions
Agentic, spec-driven variant (per sprint day)
Spec day1 spec-write, 1 execution start2approximately 25Phase 1 + Phase 2 kickoff; spec authored, execution begins
Exec day1–2 execution sessions1–265–90Phase 2 sustained; building against spec, /compact as needed
Review day1 conformance, 1 correction2approximately 20Phase 3 review; deviation notes → Phase 2 correction turn

Permission mode

Sessions start in Manual mode, where Claude asks before file writes and most shell commands. Auto mode is opt-in on Amazon Bedrock and must be pinned through managed settings using defaultMode, and Anthropic has stated an intention to make it the default on cloud platforms without committing to a date. The secd3v PROTECTED deployment guide permits auto mode with organisationally defined risky actions held as ask rules, which removes the approval fatigue that trains developers to click through everything.

Auto mode does not remove the safety gate, it moves it. Instead of prompting, it routes every tool call through a separate classifier that evaluates whether the action is irreversible, destructive or aimed outside the developer's environment, with data exfiltration in a hard-deny category the classifier is designed never to approve. Permission rules still fire before the classifier, so the ask rules implementing the ISM-2113 human approval gate over the organisationally defined risky set continue to hold in auto mode. Where the CCS GitLab MCP server is in use, its checkpoints on high-risk platform operations add a second gate covering merge approvals and mutations on the code platform, which the agent's own controls do not reach. The gateway carries no human-in-the-loop mechanism of its own.

Two cost consequences follow, and neither appears in section 04. The classifier consumes a small number of extra tokens on every tool call, and on Bedrock those calls count toward token usage, where Anthropic stopped charging the overhead on the subscription plans in August 2026. Auto mode also falls back to manual approvals after three consecutive blocks or twenty across a session, which adds approval turns to the affected session.

The session definitions above assume auto mode with ask rules. Manual mode across the board raises turn counts for the same work, because each approval round trip adds a turn, so a deployment that mandates Manual should uplift the turn counts before applying section 06.

Overall, the pattern determines the dominant cost driver: DevSecOps spends on cache reads across many short sessions, while Agentic spends on regular input inside a small number of long ones.

Section 04

Token Modelling & Fixed Context

Claude Code's context is cumulative. Every API call processes the full conversation history to date. The fixed system context is the prime candidate for prompt caching, written once per session and read cheaply on every subsequent turn.

Use case A (no MCP)
System prompt (agent instructions)3,900 tok
Built-in tool definitions15,600 tok
CLAUDE.md + memory files1,500 tok
Total fixed cached context21,000 tokens
Use case B (with GitLab MCP)
All use case A context21,000 tok
GitLab MCP tool names and schemas9,000 tok
Total fixed cached context30,000 tokens
GitLab MCP still costs more tokens, and the deferral saving does not apply. Claude Code can defer MCP tool definitions and load them on demand through tool search, which would reduce the GitLab overhead to tool names alone. That saving does not reach this deployment, for two independent reasons.
  • The deferral threshold is not reached – deferral engages only when MCP tool descriptions exceed roughly 10% of the context budget. One GitLab server at approximately 9,000 tokens sits far below that threshold on a 200,000-token window and further below it on a 1M-token window, so all schemas load upfront.
  • Tool search is not reliable on the Bedrock API – Anthropic excludes the tool search beta header for Bedrock, and an open Claude Code defect records tool search activating client-side and generating tool reference blocks that Bedrock rejects with a 400 error.
The 30,000-token figure is therefore the working figure rather than a conservative worst case, and the use case B (with GitLab MCP) cost tables in section 06 are built on it. Where deferral does engage it expands tools as inline references rather than by modifying the tools array, so it does not invalidate the cached prefix. The invalidation risk here comes from the managed MCP configuration changing between sessions, not from loading a tool within one.
The IDE connection adds per-turn context. When the CLI runs inside Visual Studio Code it connects automatically to the editor's in-process MCP server. While connected, it includes the current editor selection and the path of the active file as context on every prompt sent. The transcript shows a selected-lines notice when this happens. This is variable, uncached and unbudgeted in the totals below. A Read deny rule on a sensitive path prevents both the selected text and the open-file notice for that file from reaching the model.

Daily token volumes

↔ scroll if needed
PatternUse caseTierCache writesCache readsRegular inputOutput
DevSecOps (Brownfield): daily token totals
DevSecOpsA (no MCP)Light179,000924,000174,00028,900
Medium184,0001,218,000320,00046,200
Heavy254,0001,680,000467,50065,200
DevSecOpsB (with MCP)Light251,0001,320,000226,50033,400
Medium256,0001,740,000393,30053,100
Heavy353,0002,400,000569,60074,700
Agentic (Greenfield): daily token totals
AgenticA (no MCP)Light115,000945,000185,00030,000
Medium62,0001,008,000626,00040,000
Heavy77,0001,848,0001,914,00091,500
AgenticB (with MCP)Light160,0001,350,000207,50030,000
Medium80,0001,440,000661,00043,000
Heavy95,0002,640,0001,971,50096,000
Figures are calibrated on the previous tokenizer. For Sonnet 5 and Opus 4.8, multiply by 1.30 before pricing; section 06 applies that multiplier. Cache writes cover the fixed system context on the opening turn of each session plus the incremental writes that caching produces as the conversation grows. Cache reads are prefix re-reads on turns 2–N. Regular input is non-cached conversation context.

Session overheads outside the table above

Seven behaviours consume tokens outside the session shapes in section 03. Three of them bill a full-context turn while the developer is doing nothing. Budget for them at the tier level and control them at the gateway.

DriverMechanismControl
Compaction/compact reads the conversation it summarises, so compacting a large context is itself a large request, and the following turn writes a fresh prefix. /clear costs nothingPrefer /clear between unrelated tasks. Reserve /compact for continuity within one task
/goal conditionsA separate evaluator re-checks the goal after every turn, and idle check-ins start a new turn carrying full context while background work runsSet CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 where goals are not required
Stop hooksA Stop hook blocks the turn from ending until the check passes, for up to 8 consecutive blocksKeep the check cheap and fast. Budget up to 8 extra turns per gated task
Scheduled tasksFires on its interval and sends full context whether or not the session is activeDisable by policy unless a named use case justifies it
Cross-session messagingA message from another session arrives as a new turn with full contextSet crossSessionInbound to hold
/batchSplits a change across 5–30 subagents, each in its own worktree, each opening a merge requestTreat as an uncapped multiplier alongside agent teams. Govern at the service budget, not by guidance
Auto mode classifierA separate classifier evaluates every tool call and consumes a small number of extra tokens per call. Those calls count toward token usage on Bedrock. Auto mode also falls back to manual approvals after three consecutive blocks or twenty across a session, adding approval turnsUnavoidable while auto mode is enabled, and the safety case supports keeping it. Measure the per-tool-call overhead at the 30-day review and fold it into the tier volumes
Agent teams use approximately 7 times more tokens than standard sessions when teammates run in plan mode, because each teammate maintains its own context window and runs as a separate instance. They remain disabled by default behind CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1. Where a team approves them, four controls limit the exposure: Sonnet 5 for teammates, small teams, focused spawn prompts (teammates load CLAUDE.md, MCP servers and skills automatically), and shutting teammates down when their work is done. The service budget ceiling remains the primary control.

Overall, the fixed context is cheap to cache and expensive to ignore, and the six background drivers above are the difference between a modelled budget and an actual invoice.

Section 05

Model Access Policy & Recommended Splits Service-enforced

Opus costs 2.5 times Sonnet 5 per input token, and 3.25 times once the tokenizer difference is included. Two developer tiers define access. The secd3v Claude Code Service enforces both by role-based routing, per user, per team and organisation wide.

Standard Developer No Opus
DevSecOps: 30% Haiku 4.5 / 70% Sonnet 5
Agentic (standard): 15% Haiku 4.5 / 85% Sonnet 5
Agentic (spec-driven): 20% Haiku 4.5 / 72% Sonnet 5 / 8% Opus 4.8
DevSecOps 30 / 70
Agentic standard 15 / 85
Agentic spec-driven 20 / 72 / 8
  • Applies to the majority of an engineering organisation
  • DevSecOps Haiku share unchanged; doc updates, pipeline triage and syntax checks are genuinely Haiku-suitable
  • Spec-driven opens Opus access at Phase 1 under the spec-write session profile
  • Spec-driven Haiku share rises to 20%; Phase 3 conformance is pattern matching, not reasoning
Senior / Approved Limited Opus
DevSecOps: 25% Haiku 4.5 / 65% Sonnet 5 / 10% Opus 4.8
Agentic (standard): 10% Haiku 4.5 / 75% Sonnet 5 / 15% Opus 4.8
Agentic (spec-driven): 18% Haiku 4.5 / 67% Sonnet 5 / 15% Opus 4.8
DevSecOps 25 / 65 / 10
Agentic standard 10 / 75 / 15
Agentic spec-driven 18 / 67 / 15
  • Applies to tech leads, security engineers and principal developers
  • Spec-driven: Opus share unchanged but front-loaded into Phase 1 rather than spread across all turns
  • Opus reasoning concentrated where it has maximum downstream leverage
  • Phase 1 spec sessions need an explicit Opus permit in the service
Escalation runs across three levers, not one. Model tier is the most expensive lever and should be the last one pulled. Reach for effort level first: Sonnet 5 at high effort is the correct first escalation for a standard developer facing a hard task. Session profile is the second, since plan mode and a written spec remove more wrong-direction cost than a larger model adds. Opus is the exception rather than the default response to difficulty, and the gateway enforces that ordering rather than leaving it to judgement.

Spec authoring opens Opus access to standard developers on amortisation grounds. A single Opus 4.8 spec session amortises across 40 to 65 Sonnet 5 execution turns. The overhead recovers within the first execution session through reduced file-exploration turns and eliminated wrong-direction corrections. The service implements a spec-write session profile that permits Opus and enforces Sonnet-only for subsequent sessions in the same sprint.

The Haiku share differs between patterns for a reason grounded in task shape. In DevSecOps a substantial fraction of daily interactions, including documentation strings, pipeline status checks, dependency lookups and boilerplate, genuinely do not require Sonnet-level intelligence. In standard Agentic sessions even simple tasks involve sustained multi-turn reasoning where Haiku creates quality drag and more wrong-direction attempts. In spec-driven Agentic, Phase 3 conformance review is pattern matching against defined criteria and Haiku handles it at a third of the cost.

One constraint applies to Haiku routing. Haiku 4.5 requires 4,096 tokens per cache checkpoint against 1,024 for Sonnet 5. AWS evaluates that minimum against the combined tools, system and messages token count, so the 21,000 tokens of fixed context clear it comfortably in a normal session. The threshold constrains only a subagent invocation carrying a small prompt and no tool definitions, where nothing caches and no error is returned.

Where the current generation is better, and when to stay put

Sonnet 5 and Opus 4.8 improve on their predecessors on measurable grounds, and the recommendation to select them is not a preference for the newer label.

ComparisonWhere the current model is betterWhere it costs more
Sonnet 5 against Sonnet 4.6Anthropic positions it as its most agentic Sonnet. The rate card is 33% lower, and after the tokenizer increase the measured spend is 13.3% lower for identical work, because input, output and both cache rates all scale by the same factorNothing. Sonnet 5 costs less on the rate card and less in measured spend
Opus 4.8 against Opus 4.6Minimum cache checkpoint falls from 4,096 tokens to 1,024, which widens what can be cached in short sessions. Adaptive reasoning brings effort-level control. A new system instruction can be added partway through a conversation without invalidating the system or message cachesApproximately 30% more expensive in practice at an identical rate card, because of the tokenizer
The reason to retain Sonnet 4.6 and Opus 4.6 is migration cost, not model quality. Where an application, prompt library, evaluation suite or agent has been built and validated against a specific model, moving it to a newer model requires re-testing, and any behavioural difference requires remediation before the application can be trusted in production. That work is real, it competes with delivery, and in a high compliance environment it carries its own evidence burden. Retain 4.6 for those workloads, price them at the 4.6 rates in section 06, and schedule the migration as planned work with a testing budget rather than treating it as a configuration change. For any new workload, select Sonnet 5 or Opus 4.8.

Overall, the standard-developer split costs 18% less than the senior split on DevSecOps and 22% less on Agentic, and the difference is entirely the Opus share.

Section 06

Complete Cost Reference

Cost per developer per month (USD), 22 working days, 5-minute TTL, Australian endpoint rates, with the tokenizer multiplier applied to Sonnet 5 and Opus 4.8. Standard = no-Opus split. Senior = limited-Opus split. Per-model pure costs shown only for building custom blends.

How the figures are derived

Each monthly figure applies the section 02 rates to the section 04 daily token volumes across 22 working days, then blends the per-model results by the section 05 splits.

monthly cost = 22 days × ( writes × write rate + reads × read rate + regular input × input rate + output × output rate ) ÷ 1,000,000

Two assumptions sit behind the volumes and should be stated rather than left implicit. The opening turn of each session carries its context as a cache write rather than as regular input, which is why both the cache-read and regular-input derivations subtract the session count from the turn count. Cache writes then exceed the simple product of sessions and fixed context, because caching writes incrementally as the conversation grows, and that increment is the residual in the cache-write column. On that basis the section 04 volumes follow from the section 03 session definitions as set out below.

↔ scroll if needed
QuantityDerivationAgreement with the published volumes
Turn countsSum of turns per session type across the tier's session mixExact for all six tiers
Output, use case A (no MCP)Sum of turns multiplied by per-turn output for each session typeExact for all six tiers. DevSecOps heavy: 25 × 400, plus 36 × 700, plus 30 × 1,000 = 65,200
Cache reads, use case A (no MCP)(Turns less sessions) × 21,000Exact for all six tiers. DevSecOps heavy: 80 read-turns × 21,000 = 1,680,000
Cache reads, use case B (with GitLab MCP)(Turns less sessions) × 30,000Exact for all six tiers. Agentic heavy: 88 read-turns × 30,000 = 2,640,000
Regular input, use case A (no MCP)(Turns less sessions) × per-turn regular inputWithin 0.1–2.7%, the variance being a small allowance carried in the published volumes
Regular input and output, use case B (with GitLab MCP)Use case A volumes plus a per-turn uplift for MCP payloadsRequires the uplift below, which the session definitions do not by themselves supply
The GitLab MCP per-turn uplift is an assumption, not a derived figure. Use case B (with GitLab MCP) carries more regular input and more output than use case A (no MCP) beyond the 9,000-token difference in fixed cached context: approximately 1,010 to 1,122 additional regular-input tokens per turn on DevSecOps, 450 to 700 on Agentic, and an output uplift of up to 105 tokens per turn.

The rationale holds, because a GitLab call returns a payload that enters the conversation as an uncached tool result, whether that is a merge request diff, a pipeline log, an issue body or a review comment thread. DevSecOps sits higher because merge request review and pipeline triage call GitLab on most turns, while a long autonomous session touches GitLab less often relative to its turn count. Treat the uplift as the second-largest single input to the use case B figures after the fixed context, and verify it in the 30-day review by comparing regular input per turn on sessions that call GitLab against sessions that do not.

DevSecOps (Brownfield): use case A (no MCP)

Policy / ModelSplitLight / moMedium / moHeavy / mo
Standard developer3070$32.56$46.32$65.62
Senior (limited Opus)256510$39.78$56.59$80.17
Saving: Standard vs Senior−$7.22 (18%)−$10.27 (18%)−$14.55 (18%)
Per-model pure reference
Haiku 4.5n/a$15.36$21.85$30.95
Sonnet 5n/a$39.93$56.80$80.47
Sonnet 4.6n/a$46.08$65.54$92.86
Opus 4.8n/a$99.83$142.01$201.19
Opus 4.6n/a$76.79$109.24$154.76

DevSecOps (Brownfield): use case B (with GitLab MCP)

Policy / ModelSplitLight / moMedium / moHeavy / mo
Standard developer3070$43.06$59.14$83.34
Senior (limited Opus)256510$52.60$72.26$101.81
Saving: Standard vs Senior−$9.55 (18%)−$13.11 (18%)−$18.48 (18%)
Per-model pure reference
Haiku 4.5n/a$20.31$27.90$39.31
Sonnet 5n/a$52.81$72.53$102.20
Sonnet 4.6n/a$60.93$83.69$117.93
Opus 4.8n/a$132.01$181.34$255.51
Opus 4.6n/a$101.55$139.49$196.55

Agentic (Greenfield): use case A (no MCP)

Policy / ModelSplitLight / moMedium / moHeavy / mo
Standard developer1585$32.74$57.36$151.49
Standard developer, spec-driven est.20728$30.39$44.15$108.76
Saving: Standard vs Spec-Driven−$2.35 (7%)−$13.21 (23%)−$42.73 (28%)
Senior (limited Opus)107515$41.96$73.52$194.18
Senior, spec-driven est.186715$33.96$49.34$121.55
Saving: Standard vs Senior−$9.23 (22%)−$16.16 (22%)−$42.69 (22%)
Per-model pure reference
Haiku 4.5n/a$13.87$24.30$64.19
Sonnet 5n/a$36.07$63.19$166.90
Sonnet 4.6n/a$41.62$72.91$192.58
Opus 4.8n/a$90.17$157.98$417.25
Opus 4.6n/a$69.36$121.52$320.96

Agentic (Greenfield): use case B (with GitLab MCP)

Policy / ModelSplitLight / moMedium / moHeavy / mo
Standard developer1585$39.55$63.97$161.87
Standard developer, spec-driven est.20728$37.19$50.35$118.42
Saving: Standard vs Spec-Driven−$2.36 (6%)−$13.62 (21%)−$43.45 (27%)
Senior (limited Opus)107515$50.69$81.99$207.48
Senior, spec-driven est.186715$41.57$56.28$132.36
Saving: Standard vs Senior−$11.14 (22%)−$18.02 (22%)−$45.61 (22%)
Per-model pure reference
Haiku 4.5n/a$16.76$27.10$68.59
Sonnet 5n/a$43.57$70.47$178.33
Sonnet 4.6n/a$50.28$80.47$205.77
Opus 4.8n/a$108.93$176.18$445.83
Opus 4.6n/a$83.79$135.52$342.94
Spec-driven figures are marked estimated. They model Phase 2 regular input at 13,000 tokens/turn against 25,000 for standard agentic, include the Phase 1 Opus overhead at one spec session per five execution sessions, and assume Phase 3 conformance sessions run on Haiku. Actual results depend on spec quality and sprint cadence.

Pattern comparison: standard developer policy

PatternUse caseLight / moMedium / moHeavy / moDominant cost driver at heavy
DevSecOpsA (no MCP)$32.56$46.32$65.62Cache reads across 11 short sessions
AgenticA (no MCP)$32.74$57.36$151.49Regular input at 1,914k tok/day, 4.1 times DevSecOps
Spec-Driven est.A (no MCP)$30.39$44.15$108.76Regular input cut 48% by the spec; Phase 3 on Haiku
DevSecOpsB (with MCP)$43.06$59.14$83.34MCP adds 32% at light and 27% at heavy
AgenticB (with MCP)$39.55$63.97$161.87Heavy agentic runs 94% above DevSecOps heavy
Spec-Driven est.B (with MCP)$37.19$50.35$118.42Spec-driven narrows the gap against DevSecOps to 42% at heavy

Scope of these figures, and the complementary activity allowance

Every figure in this document models software development: reading and writing code, reviewing changes, generating tests, auditing for security defects, and building services against a specification. The session definitions in section 03 are drawn from that work and nothing else.

Developers also use the agent for content activities that sit alongside development and are not costed above. These include long-form documentation, release notes and changelogs, merge request and release summarisation, architecture and sequence diagrams, commit message drafting, incident write-ups, runbooks, onboarding material, and open-ended explanations of unfamiliar code. Some of this work appears in the model only incidentally, through the documentation update and quick explanation examples in the DevSecOps micro session, and the rest is absent.

These sessions carry a different token shape from development work. They run short, they read a bounded amount of context, and they produce a great deal of output, which is the most expensive token category at five times the input rate. One such session per day, modelled at 6 turns with 4,000 tokens of regular input per turn and 1,500 tokens of output per turn on the standard developer split, costs approximately $5.22 per month.

Base figure, standard policyOne session / dayTwo sessions / day
DevSecOps light, use case A (no MCP), $32.5616.0%32.1%
DevSecOps heavy, use case A (no MCP), $65.628.0%15.9%
DevSecOps heavy, use case B (with GitLab MCP), $83.346.3%12.5%
Agentic heavy, use case A (no MCP), $151.493.4%6.9%
Agentic heavy, use case B (with GitLab MCP), $161.873.2%6.4%
Add 10% to every figure in this section as a planning allowance. That corresponds to one to two complementary sessions per developer per day and sits mid-range across the patterns. Use 15% where merge request summarisation, release notes or documentation form a stated part of the role, and for developers at the light DevSecOps tier, where the allowance is proportionally larger against a smaller base. Agentic tiers are less sensitive, because a heavy autonomous session dwarfs a short content session, so 5% is defensible for a purely agentic team.
Pattern and tier, standard policyBase / moWith 10% allowance / mo
DevSecOps heavy, use case A (no MCP)$65.62$72.18
DevSecOps heavy, use case B (with GitLab MCP)$83.34$91.67
Agentic heavy, use case A (no MCP)$151.49$166.64
Agentic heavy, use case B (with GitLab MCP)$161.87$178.06
Spec-driven heavy, use case A (no MCP) est.$108.76$119.64
Spec-driven heavy, use case B (with GitLab MCP) est.$118.42$130.26
The allowance is a planning figure rather than a measured one. Replace it at the 30-day telemetry review by separating sessions whose output-to-input ratio exceeds roughly 1 to 3, which is the signature of content work rather than code work.

Overall, spec-driven development delivers the largest single return on investment available in the model, cutting heavy agentic spend by 28% for the cost of one authoring session per sprint, and a 10% allowance covers the content activities these figures exclude.

Section 07

Prompt Caching & TTL Selection

Prompt caching is the largest cost lever in this model. Without it, the fixed system context bills as full-price regular input on every API call. Prompt caching is enabled by default on the Bedrock API, and the 5-minute TTL applies unless a request specifies otherwise.

38–63%Cost reduction from caching
vs no-cache baseline
>90%Measured cache hit rate
on the 5-min default
$3–$14Monthly cost of selecting 1-hr TTL
where expiry is not observed

Caching against a no-cache baseline

Standard developer blend, 5-minute TTL, AU regional.

↔ scroll if needed
PatternUse caseTierNo cache / mo5-min TTL / moSaving
DevSecOps
DevSecOpsA (no MCP)Light$72.93$32.56$40.37 (55%)
Medium$100.20$46.32$53.88 (54%)
Heavy$139.93$65.62$74.31 (53%)
DevSecOpsB (with MCP)Light$100.79$43.06$57.73 (57%)
Medium$136.20$59.14$77.06 (57%)
Heavy$189.62$83.34$106.29 (56%)
Agentic
AgenticA (no MCP)Light$79.67$32.74$46.93 (59%)
Medium$108.28$57.36$50.93 (47%)
Heavy$245.38$151.49$93.89 (38%)
AgenticB (with MCP)Light$106.66$39.55$67.11 (63%)
Medium$136.84$63.97$72.87 (53%)
Heavy$296.21$161.87$134.34 (45%)

The 5-minute TTL is the correct default, and the evidence is the vendors' own

AWS and Anthropic both publish the same guidance, and both frame the 5-minute cache as the default for regularly used prompts rather than as a limitation to be worked around.

  • The TTL resets on every hit, at no charge – AWS states that the cache has a time to live which resets with each successful cache hit, that the context is preserved during that period, and that the cache expires only if no cache hits occur within the window. Anthropic states the same mechanism, that the cache is refreshed for no additional cost each time the cached content is used.
  • AWS recommends the 5-minute cache for prompts used more often than every five minutes – the Bedrock user guide's best practice for Anthropic models states that where prompts are used at a regular cadence, meaning more frequently than every five minutes, you should continue to use the 5-minute cache, since it will continue to be refreshed at no additional charge.
  • AWS reserves the 1-hour cache for three named scenarios – prompts used less frequently than every five minutes but more frequently than every hour, cases where latency matters and follow-up prompts may arrive beyond five minutes, and improving rate limit use because cache hits are not deducted against the rate limit. An interactive Claude Code session is none of the three.

A Claude Code session is a tool-use loop in which one conversational turn issues many API requests seconds apart, so the gaps that matter are genuine idle periods rather than turn boundaries. That is precisely the regular cadence the vendor guidance describes. Production telemetry agrees: cache hit rates exceed 90% on the 5-minute TTL across Claude Code and Claude Code Service usage.

One mechanical detail bounds the argument. Anthropic states that the lifetime is measured from the start of the request that writes or reads the cache entry rather than from the end of its response, and that generation time counts against the lifetime, so a response taking four minutes to stream leaves roughly one minute for the follow-up request. Long, high-effort Opus turns are the one interactive case that can expire a 5-minute cache without the developer pausing at all.
break-even: 1-hr write = 2.0 × base input · 5-min write = 1.25 × base input · cache read = 0.1 × base input premium = 60% on every cache-write token, not only the opening prefix, because caching writes incrementally on almost every turn

What the 1-hour TTL costs when expiry does not occur

↔ scroll if needed
PatternUse caseTier5-min TTL / mo1-hr TTL / mo1-hr premium
DevSecOps: the premium is largest, across eleven short sessions a day
DevSecOpsA (no MCP)Light$32.56$39.45+$6.89
Medium$46.32$53.40+$7.08
Heavy$65.62$75.39+$9.77
DevSecOpsB (with MCP)Light$43.06$52.71+$9.66
Medium$59.14$68.99+$9.85
Heavy$83.34$96.92+$13.58
Agentic: close to indifferent either way
AgenticA (no MCP)Light$32.74$37.67+$4.93
Medium$57.36$60.01+$2.66
Heavy$151.49$154.79+$3.30
AgenticB (with MCP)Light$39.55$46.40+$6.85
Medium$63.97$67.39+$3.43
Heavy$161.87$165.94+$4.07
The two patterns sit at opposite ends of that table. DevSecOps pays the premium eleven times a day across eleven short sessions, so a 1-hour profile costs a heavy DevSecOps developer $10 to $14 per month for a benefit the telemetry does not show. Heavy agentic pays it twice a day and sits close to indifferent at $3 to $4 per month.

When to select the 1-hour TTL

5-minute TTL: default for everything
  • Requires no configuration; it is also the Bedrock default
  • All interactive DevSecOps sessions, micro through extended
  • All standard agentic sessions
  • All spec-driven Phase 2 execution sessions
  • Fully automated CI/CD pipelines and sequential claude -p invocations
  • Pre-commit hook automation running in-process
1-hour TTL: per-profile exception, on evidence
  • Long single turns. The lifetime measures from the start of the request, not the end of the response, and generation time counts against it. A four-minute stream leaves about one minute for the follow-up. Opus 4.8 at high effort is the exposure
  • Human review gaps. MR review, compliance reading and multi-file approval, particularly in Manual permission mode
  • CI pipeline waits. GitLab MCP workflows that block on a pipeline
  • Cross-session cache grouping. Impossible on a 5-minute TTL. Bedrock uses organisation-level cache isolation, so sharing is available. Measure the benefit before claiming it
  • Rate limit headroom. Cache hits are not deducted against rate limits, so fewer re-writes buys throughput on a constrained TPM quota
ENABLE_PROMPT_CACHING_1H is a process-level environment variable, not a per-request setting. It applies to every session a developer runs and cannot be selected per session type, so this is a per-developer-profile decision.

Cache misses a longer TTL will not fix

Where residual misses come from invalidation rather than expiry, a longer TTL buys nothing and costs the premium. Six documented causes apply.

CauseEffectRemedy
Effort level changed mid-sessionThe resolved effort value renders into the prompt, so a change invalidates message blocksSet effort once per session and hold it. See section 08
Thinking configuration changedSame mechanism as effortSet once per session
Tools section changedCheckpoints chain in the order tools, system, messages, so modifying tools invalidates the system and messages caches with itHold the managed MCP configuration stable between sessions
Image addedAdding an image anywhere in the prompt invalidates message blocksExpect a full re-write after a pasted screenshot
20-block lookback exceededAutomatic prefix checking looks back approximately 20 content blocks from the checkpoint, and static content beyond that range is not foundOccurs on long parallel tool sequences. Additional checkpoints are the documented remedy, up to the four-checkpoint maximum
Cross-region routing under loadAt times of high demand, cross-region inference optimisations may lead to increased cache writesPrefer direct in-region routing for cache-sensitive workloads where capacity allows
Attribute the residual misses before changing any TTL setting. A miss caused by expiry shows a full-prefix write after an idle gap. A miss caused by invalidation shows a full-prefix write with no preceding gap, and the remedy there is configuration discipline rather than a longer TTL. /compact also produces a full cache write on the following turn, because summarisation replaces the conversation prefix; that is a cost of compaction rather than a cache fault and section 04 carries it.

Overall, caching delivers 38–63% against an uncached baseline, and the 5-minute default captures that saving with no configuration and no premium.

Section 08

Developer Cost Optimisation Factors

After model policy and caching configuration, developer behaviour controls cost. Each factor below is validated against Anthropic's Claude Code best practices documentation.

🎯
Prompt specificity and context front-loading
Highest impact · All patterns
"The more precise your instructions, the fewer corrections you'll need. Reference specific files, mention constraints, and point to example patterns." A prompt like "review @auth.py lines 42–89 for SQL injection, here is the schema" costs three to five times less than "check my auth code for security issues", because Claude does not explore files to find context it was never given.

In DevSecOps, specificity keeps sessions inside their session-type scope and prevents drift from Micro into Extended territory. In agentic work, vague prompts trigger broad file scanning and each file read compounds into every subsequent turn. In spec-driven execution the spec is the specificity mechanism, but the execution prompt must still reference specific spec sections.
DevSecOpsAgenticthree to five times less regular input
📋
Plan mode before execution
Highest impact · Agentic
Enter plan mode by pressing Shift+Tab until the status bar shows plan mode on, or start the session with claude --permission-mode plan. Press Ctrl+G to open the plan in a text editor and edit it before Claude proceeds.

Plan mode helps most where the approach is uncertain, the change touches multiple files, or the developer is unfamiliar with the code. Skip it for small, clearly scoped tasks: if the diff fits in one sentence, skip the plan. Correcting direction at the plan stage costs approximately 500 tokens. Correcting after twenty turns of wrong implementation costs tens of thousands.
AgenticDevSecOpsprevents expensive wrong-direction sessions
✅
Verification targets
High impact · All patterns
Claude stops when the work looks done. Without a check it can run, "looks done" is the only available signal and the developer then becomes the verification loop, which is the most expensive loop in the workflow. Give Claude something that returns pass or fail: a test suite, a build exit code, a linter, or a script that diffs output against a fixture. The loop then closes on its own.

Ask for evidence rather than assertion: the test output, the command run and what it returned. Reviewing evidence costs less than re-running the verification. Where a task needs a harder gate, a /goal condition or a Stop hook enforces it, and section 04 carries the token cost of both.
DevSecOpsAgenticcloses the loop without the developer
📐
Spec-driven development: Phase 1 spec authoring
Highest impact · Agentic only · approximately 28% monthly saving at heavy
Author a structured specification before any agentic execution begins. Claude then executes against the spec rather than discovering scope through open-ended file exploration, the dominant cost driver in standard agentic sessions. Regular input falls from approximately 25,000 tokens per turn to approximately 13,000, and wrong-direction turns fall from 4–8 per session to 0–1.

CLAUDE.md rule. The spec belongs in a separate file referenced at session start, never embedded in CLAUDE.md. A 600-line spec in CLAUDE.md adds approximately 9,000 tokens of cache-read cost per call with no benefit. Target CLAUDE.md under 200 lines regardless.

Session sequence. Write the spec to a file, end the authoring session, then start a fresh session to execute. Run /compact at 80% context fill during Phase 2.
Ph1 Opus 4.8Ph2 Sonnet 5Ph3 Haiku 4.5approximately 48% regular input reduction
🧹
Session hygiene: /clear, /compact and /btw
High impact · DevSecOps especially
"Run /clear between unrelated tasks to reset context. If you've corrected Claude more than twice on the same issue in one session, the context is cluttered with failed approaches." A clean session with a better prompt almost always outperforms a long session with accumulated corrections. Use /rename before clearing, then claude --resume to return.

For long agentic sessions, /compact summarises history rather than clearing it, and takes focus instructions such as /compact focus on the API changes.

Two cheaper alternatives are available. /btw asks a side question whose answer never enters conversation history. The rewind menu offers summarise-from-here and summarise-up-to-here, which condense part of the conversation while leaving the rest intact.
/clear between every task/compact at approximately 80%
🔑
Model selection and effort levels
High impact · DevSecOps especially
Defaulting to Sonnet for everything is the most common unnecessary cost. Documentation strings, pipeline triage, dependency lookups, simple formatting and boilerplate scaffolding are genuinely Haiku-suitable at a third of the cost.

Set effort once per session and hold it. The resolved effort value renders into the prompt, so changing it between requests invalidates the message-block cache and forces a re-write. Choose the effort level with the session type at the start, then steer individual turns through prompt wording.

DevSecOps micro and standard: low or medium. DevSecOps extended and agentic execution: medium. Spec-driven Phase 3: low. Phase 1 authoring and novel reasoning: high.
25–30% Haiku achievable10–15% Haiku realistic
🔒
Extended thinking: model-dependent controls
High cost when unmanaged
Thinking tokens bill as output tokens, the most expensive category, and Anthropic enables extended thinking by default because it materially improves complex planning and reasoning.

The control depends on the model generation. On Sonnet 5 and Opus 4.8 only adaptive reasoning exists; effort level is the control and these models ignore a nonzero MAX_THINKING_TOKENS. On Sonnet 4.6 and Opus 4.6, MAX_THINKING_TOKENS applies in combination with CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1.

On Sonnet 5 and Opus 4.8, thinking display defaults to omitted rather than summarised. Telemetry that estimates reasoning cost from returned content will under-report on these models and should use the cache and token counts the Bedrock response reports instead.
effort is the control/effort for high-complexity tasks
🔍
Subagents for research and cost routing
High impact · Agentic
"Delegate research with 'use subagents to investigate X'. They explore in a separate context, keeping your main conversation clean for implementation." Subagents also serve as cost routers: specify model: haiku in the subagent configuration for file scanning, documentation lookup and log analysis, and the main Sonnet 5 session receives a summary without paying Sonnet prices for the exploration.

A verification subagent reviewing the diff in a fresh context is the cheapest quality gate available, because it sees only the diff and the criteria rather than the reasoning that produced the change. Tell the reviewer to flag only gaps affecting correctness or the stated requirements, since a reviewer asked to find gaps will find them and chasing every one produces over-engineering.
context preservation + Haiku routinglarge SAST result triage
⚙️
Preprocessing hooks
High impact · DevSecOps
A hook preprocesses data before Claude sees it. Instead of Claude reading a 10,000-line log file to find errors, a PreToolUse hook greps for ERROR and returns only the matching lines, cutting context from tens of thousands of tokens to hundreds. For a brownfield estate with verbose CI output this is the largest untapped lever in the DevSecOps pattern.

Hooks are deterministic where CLAUDE.md instructions are advisory, so they also suit any control that must happen every time.
DevSecOpstens of thousands of tokens → hundreds
🗂️
CLAUDE.md, skills and .claudeignore discipline
Medium impact · Both patterns
Keep CLAUDE.md specific, concise and under 200 lines. It loads into every session and consumes tokens on every turn. Move specialist content to Skills, which load on demand. For each line ask whether removing it would cause a mistake, and cut it if not. A bloated CLAUDE.md also causes Claude to ignore the instructions that matter.

Configure .claudeignore on day one of any brownfield project. A single accidental glob-all read of a large repository consumes 50,000–150,000 tokens in one call, most of a day's DevSecOps budget.

One failure mode is specific to spec-driven projects. Developers embed verbose spec prose into CLAUDE.md. A 600-line spec adds approximately 9,000 tokens of cache-read cost per call with no benefit. CLAUDE.md above 3,000 tokens on a spec-driven project indicates this failure is active.
DevSecOpsAgenticone bad file read can cost a day's budget
🔌
CLI tools and code intelligence plugins
Medium impact · Both patterns
CLI tools remain the most context-efficient way to reach an external service, because they add no per-tool listing to the context window. Prefer them over an MCP server where both exist, which matters more in this deployment because the GitLab MCP schemas load upfront rather than on demand.

For typed languages, a code intelligence plugin gives precise symbol navigation instead of text search. One go-to-definition call replaces a grep followed by reads of several candidate files, and the installed language server reports type errors automatically after edits so Claude catches mistakes without running a compiler.
DevSecOpsAgenticverify with /context in session
🛠️
Adding claude to PATH
Prerequisite, not an optimisation
Running claude in the integrated terminal requires the standalone CLI install on the shell PATH. Installing the graphical extension does not provide it, because the extension bundles a private copy of the CLI for its own chat panel.

Without the PATH entry, claude -p (pre-commit automation and fan-out loops), claude --permission-mode plan, claude --resume, claude mcp list (the managed MCP validation check) and claude --version are all unavailable, and each is a cost-reducing behaviour elsewhere in this document.

One limit applies to the automation cases. The gateway records machine-to-machine access as an open gap, and a CI job reaches the CCS GitLab MCP server tool surface but not the gateway's inference path. Non-interactive inference is therefore available under a developer's own session on their workstation, and a pipeline that needs inference is not currently served. No figure in section 06 depends on CI inference, because every tier models interactive developer sessions.

Install the CLI to the managed, administrator-owned path the deployment guides already mandate, and add that path to the system PATH through the managed environment configuration rather than the user profile. The application control position does not change. /clear and /compact are in-session commands and do not depend on PATH, but starting the CLI does.
DevSecOpsAgenticverify: claude --version · claude mcp list
📊
Token telemetry: measure first
Foundation for all other optimisation
The secd3v Claude Code Service prices and attributes every request, aggregated per user, per team, per model and organisation wide. Claude Code calls Bedrock directly and sends no usage metrics to Anthropic, so the service is the source of attribution. Section 11 lists the signals worth alerting on and the 30-day review cadence.
DevSecOpsAgentic30-day review cadence

Overall, prompt specificity, plan mode and spec authoring carry the highest return of the twelve factors, and the remainder protect that return rather than add to it.

Section 09

Opus Approval Guide Senior Access

Opus is justified where the task requires novel reasoning under genuine ambiguity, sustained autonomous operation over many turns, or where downstream error costs are high enough that a reasoning gap materially changes outcomes. Source the current benchmark comparison from Anthropic's model cards before quoting figures in a business case.

Spec authoring introduces an Opus justification that applies to all developer tiers. A single Opus 4.8 spec session producing a tight 1,500–2,500 token specification amortises across 40–65 Sonnet 5 execution turns, and the cost recovers within the first execution session.

Task or scenario
Pattern
Opus justified?
Rationale
Spec authoring for a greenfield service or controlled migration sprint
Agentic
Yes, all tiers
Phase 1 only. Amortised across 40–65 Sonnet 5 execution turns. Architectural scope and interface decisions are where the reasoning gap is most material
Security vulnerability exploit chain assessment, novel threat vectors
DevSecOps
Yes
Novel reasoning under genuine ambiguity
Compliance gap analysis against complex regulatory controls
DevSecOps
Yes
Multi-control reasoning where errors carry significant downstream cost
Architectural design for a new greenfield service
Agentic
Yes
A single high-value planning session prevents many expensive wrong-direction turns
Multi-service brownfield refactor with ambiguous legacy coupling
Agentic
Conditional
Use Sonnet 5 at high effort in plan mode first. Escalate only if plan proposals are inadequate twice
Security threat modelling, genuinely novel scenarios
DevSecOps
Conditional
Opus for novel scenarios. Sonnet 5 handles known patterns including OWASP Top 10 and CVE triage
Spec execution, implementing against an authored spec
Agentic
No
Phase 2. The spec provides the reasoning frame. Haiku is eligible for bounded implementation turns
Standard MR code review, 1–5 files
DevSecOps
No
Routine review is Sonnet 5 territory
Feature implementation, well-defined greenfield module
Agentic
No
Well-defined implementation is Sonnet 5 territory. Plan mode compensates for the narrower reasoning
Conformance review, verifying output against a spec
Agentic
No
Phase 3. Pattern matching against defined criteria, which is a Haiku task
Documentation, comments and type annotations
DevSecOps
No
Haiku task. Opus costs five times the input for equivalent output
Pipeline failure triage and CI configuration
DevSecOps
No
Pattern matching, not novel reasoning
Automated CI/CD pipeline tasks with no human in the loop
Agentic
No
No human verifies the reasoning, and Sonnet 5 gives the better cost-reliability trade

Overall, four scenarios justify Opus outright, two are conditional on Sonnet 5 failing first, and seven do not justify it at any tier.

Section 10

Mandatory Service Configuration

↔ scroll if needed
SettingRequirementConsequence if omitted
Caching
ENABLE_PROMPT_CACHING_1HLeave unset so the Bedrock 5-minute default applies. Enable per developer profile only on measured expiry, per section 07Enabling it estate-wide costs $3–$14 per developer per month for no measured benefit
Background token consumption
CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0Set unless goal-driven sessions are an approved workflowIdle check-ins send full context on an interval
crossSessionInbound = holdPrevents inbound cross-session messages arriving as full-context turnsUncontrolled turns on idle sessions
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMSLeave unset unless approved per team, with a service budget ceilingApproximately 7 times token usage
Endpoint
Standalone CLI on the managed system PATHPer section 08Automation, plan mode startup and MCP validation all unavailable
managed-mcp.jsonPer the deployment guides, listing only the CCS GitLab MCP server, and held stable between sessionsUncontrolled MCP context and tool authority, plus cache invalidation when the tools section changes

Overall, the agent teams flag carries the largest single financial risk at roughly 7 times token usage, and the caching default carries the largest recurring one.

Section 11

Telemetry & Cost Attribution

Claude Code sends no usage metrics to Anthropic, so attribution comes from the request path. Because every inference request passes through the CCS gateway, the gateway is the natural and complete point of measurement, and the accounting planes described in the CCS gateway high level design are where per-user consumption is recorded. The gateway also assumes its upstream role per user with session tags, which carries the same attribution into the AWS audit trail, so CloudTrail corroborates the gateway record rather than substituting for it.

Telemetry signals

Every signal below derives from request metadata recorded at the gateway, apart from the CLAUDE.md check, which is endpoint-side.

SignalIndicates
Caching
Cache hit rate below 70%Poor session hygiene, or one of the invalidation causes in section 07
Full-prefix cache write after an idle gapTTL expiry. The only evidence that justifies enabling the 1-hour profile for that developer
Full-prefix cache write with no preceding gapInvalidation rather than expiry. Check for effort changes, tools-section changes or pasted images
Cache write volume rising with no change in session patternRouting behaviour under load where a multi-region profile is in use, per section 02
Cache token counts at zeroA prompt below the minimum cacheable length for the model in use
Cache write records reporting a 1-hour TTLThe 1-hour setting is active on a profile that may not warrant it
Context discipline
Regular input above 15,000 tok/turn on DevSecOpsBroad prompting, or a missing /clear
Regular input above 18,000 tok/turn during spec-driven Phase 2The spec is not being referenced and Claude is still file-exploring
CLAUDE.md above 3,000 tokens on a spec-driven projectSpec prose embedded in CLAUDE.md. Endpoint-side check only, because the gateway records metadata and cannot inspect the file
Model policy
Haiku share below 15% on DevSecOpsModel discipline not applied

Overall, customers should conduct a recurring 30-day telemetry review against the signals above, restating the estimated inputs in this model from measured data and reissuing the affected tables.