Enterprise cost considerations for DevSecOps (Brownfield), Agentic (Greenfield) and Spec-Driven development using the Claude Code CLI with the secd3v Claude Code Service, with or without GitLab via the secd3v GitLab MCP service. Covers the AU rate card and cache mechanics, the model access policy, prompt caching and TTL selection, developer optimisation, mandatory service configuration, and how spec-driven workflows modify token consumption, session shape and model allocation within the Agentic pattern. Every inference request reaches Amazon Bedrock in ap-southeast-2 (Sydney) and ap-southeast-4 (Melbourne) through the CCS gateway.
secd3v Claude Code Service usage in high compliance environments organises into two foundational patterns, DevSecOps (Brownfield) and Agentic (Greenfield), with a third mode, Spec-Driven Development, that modifies how Agentic sessions run. All three operate across Haiku 4.5, Sonnet 5 and Opus 4.8, with Sonnet 4.6 and Opus 4.6 retained for validated legacy workloads. Every cost carries the 10% premium that Australian endpoints attract, which is the price of keeping inference onshore and is not optional for a customer operating to the ISM PROTECTED baseline.
Prompt caching is the largest single lever, cutting cost by 38 to 63% against an uncached baseline. TTL selection within it is a second-order decision that favours the 5-minute default, and section 07 sets out the evidence. Spec-driven development adds a structural second lever. It front-loads reasoning into a stable cached specification, cuts regular input per execution turn by approximately 48%, and removes the most expensive tail-cost scenario in the model, the wrong-direction agentic session.
Overall, a heavy agentic developer costs $151 per month and a heavy DevSecOps developer $66 before that allowance, provided the service enforces the configuration in section 10.
secd3v pins all inference to Australian Amazon Bedrock endpoints in ap-southeast-2 (Sydney) and ap-southeast-4 (Melbourne), reached through the AU cross-region inference profile or, in Melbourne, through direct in-region routing. Both options carry a 10% premium over the global endpoint, and that premium is included in every figure in this document. Anthropic's Bedrock documentation states the premium and names the AU inference profile as the mechanism for routing across Australian regions within the geography.
Per million tokens (MTok), on-demand rates, Australian endpoints. Rates derive from Anthropic's published rate card with the 10% Australian endpoint premium applied. Confirm the effective rate in the AWS console, or through the AWS Price List API, before using these figures for contracted forecasting.
| Model | Role | Input | Output | 5-min write | 1-hr write | Cache read |
|---|---|---|---|---|---|---|
| Haiku 4.5 | Routine tasks, subagents, conformance review | $1.10 | $5.50 | $1.375 | $2.20 | $0.11 |
| Sonnet 5 | Primary workhorse | $2.20 | $11.00 | $2.75 | $4.40 | $0.22 |
| Sonnet 4.6 | Validated legacy workloads only | $3.30 | $16.50 | $4.125 | $6.60 | $0.33 |
| Opus 4.8 | Maximum intelligence, senior-approved | $5.50 | $27.50 | $6.875 | $11.00 | $0.55 |
| Opus 4.6 | Validated legacy workloads only | $5.50 | $27.50 | $6.875 | $11.00 | $0.55 |
AWS publishes the operative limits for the Bedrock API, and all five models in the catalogue support both TTL options with up to four cache checkpoints per request.
| Model | Min tokens per checkpoint | Max checkpoints | Supported TTL |
|---|---|---|---|
| Haiku 4.5 | 4,096 | 4 | 5 minutes, 1 hour |
| Sonnet 5 | 1,024 | 4 | 5 minutes, 1 hour |
| Sonnet 4.6 | 1,024 | 4 | 5 minutes, 1 hour |
| Opus 4.8 | 1,024 | 4 | 5 minutes, 1 hour |
| Opus 4.6 | 4,096 | 4 | 5 minutes, 1 hour |
tools, system and messages sections rather than each section individually. With 21,000 tokens of fixed context in every session, the 4,096-token minimum on Haiku 4.5 and Opus 4.6 is comfortably exceeded, so the threshold constrains only a subagent invocation carrying a genuinely small prompt and no tool definitions.tools, then system, then messages. Modifying the tools section invalidates the system and messages caches with it.Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, and Sonnet 4.6 and earlier use the previous one. In the current catalogue this separates Sonnet 5 and Opus 4.8 from Sonnet 4.6, Opus 4.6 and Haiku 4.5. Rate cards do not reflect it, so the effect lands entirely on token volume and stays invisible in a price comparison.
| Comparison | Rate difference | Token difference | Net effect on spend |
|---|---|---|---|
| Sonnet 5 against Sonnet 4.6 | $2.20/$11.00 against $3.30/$16.50 | approx. +30% | Sonnet 5 approx. 13% cheaper |
| Opus 4.8 against Opus 4.6 | Identical at $5.50/$27.50 | approx. +30% | Opus 4.8 approx. 30% more expensive |
Both Australian regions serve the catalogue, and they differ in how a request reaches them. Sydney routes through the AU cross-region inference profile. Melbourne supports the AU profile and also direct single-region routing without a profile. Both attract the 10% premium.
Sonnet 5, Opus 4.8 and Sonnet 4.6 carry a 1M token context window at standard rates with no long-context surcharge. A larger window delays auto-compaction, so sessions accumulate more context before summarisation and per-turn regular input rises even though the per-token rate does not. Treat this as a token-volume effect rather than a pricing effect.
The default Bedrock quota is 2 million input tokens per minute, extensible to 4 million on request without additional Anthropic approval, with requests-per-minute limits enforced separately by AWS. Cache hits are not deducted against rate limits, which makes caching a throughput lever as well as a cost lever.
Overall, the Australian rate card carries a 10% sovereignty premium and a 30% tokenizer effect on the two current-generation models, and cross-region routing can raise cache-write volume independently of developer behaviour.
Brownfield development means working on an existing production system with live users, an established architecture, years of accumulated code, technical debt and security constraints under active enforcement. A developer cannot start fresh, cannot discard a wrong implementation without consequence, and cannot hold broad autonomous permissions without risk. In government, defence and high compliance contexts this describes the overwhelming majority of daily work.
DevSecOps is the direct response to those constraints rather than a methodology that happens to sit alongside them. Where a codebase has a security posture that cannot be accidentally degraded and a production environment where mistakes have immediate consequences, the human-gated, incremental, review-at-every-step workflow becomes mandatory rather than optional discipline.
This shapes how the CLI gets used. The developer says "look at auth.py lines 42 to 89 and identify any SQL injection risk" rather than "explore this codebase and make the changes you think are needed". Sessions run short, targeted and bounded, because the work demands precision over autonomy.
Greenfield development builds a new service, module or application, with no live users to disrupt, no established architecture to break and no accumulated constraints to misunderstand. A wrong-direction implementation gets discarded cheaply, so the cost of an incorrect autonomous attempt stays low. Claude explores, plans and implements, reading many files and building its own context map over long sessions.
Agentic development also suits specific brownfield scenarios, including large-scale migration sprints, comprehensive test generation and automated documentation, provided the developer enters plan mode before any execution, works in a git worktree for isolation, and treats every checkpoint as a safety gate. It does not suit routine brownfield maintenance, security operations, or any work where an incorrect autonomous change reaches production. The session risk profile determines the fit, not the label.
Spec-driven development overlays the Agentic pattern. The developer authors a structured specification covering interface contracts, data shapes, acceptance criteria, file layout, security constraints and test coverage, then executes against the spec rather than exploring freely.
SDD front-loads the reasoning cost. Reasoning concentrates in a short spec-writing session and every subsequent execution session becomes cheaper and more predictable. Wrong-direction errors surface at spec review, costing roughly 500 tokens to correct, rather than after twenty agentic turns. Regular input per execution turn falls by approximately 48%.
Spec-driven development matters in high-compliance, government and defence work because it establishes the specification as the authoritative source of truth and treats code as a derivative artefact. It replaces loose, non-deterministic prompting with unambiguous, executable contracts grounded in architectural and security constraints. It supports the traceability the Australian Information Security Manual (ISM) and ISO 27001 require, validating every change against documented requirements through automated gates in CI/CD. Security corrections propagate across future regeneration cycles, which mitigates architectural drift and preserves individual human accountability for AI-enabled outcomes in mission-critical environments.
| Tier | Session mix / day | Sessions | Turns | Developer profile |
|---|---|---|---|---|
| DevSecOps | ||||
| Light | 5 micro, 3 standard | 8 | 52 | Part-time AI assistance; quick reviews and consultations |
| Medium | 3 micro, 4 standard, 1 extended | 8 | 66 | Active developer; daily code review, security checks, bug fixes |
| Heavy | 5 micro, 4 standard, 2 extended | 11 | 91 | Lead developer / security engineer; audits, MR reviews, compliance |
| Agentic, standard | ||||
| Light | 5 light sessions | 5 | 50 | Light autonomous tasks; feature additions, focused builds |
| Medium | 2 medium sessions | 2 | 50 | Active builder; feature implementation, module construction |
| Heavy | 1 heavy session, 1 medium session | 2 | 90 | Power developer; new service construction, long autonomous sessions |
| Agentic, spec-driven variant (per sprint day) | ||||
| Spec day | 1 spec-write, 1 execution start | 2 | approximately 25 | Phase 1 + Phase 2 kickoff; spec authored, execution begins |
| Exec day | 1–2 execution sessions | 1–2 | 65–90 | Phase 2 sustained; building against spec, /compact as needed |
| Review day | 1 conformance, 1 correction | 2 | approximately 20 | Phase 3 review; deviation notes → Phase 2 correction turn |
defaultMode, and Anthropic has stated an intention to make it the default on cloud platforms without committing to a date. The secd3v PROTECTED deployment guide permits auto mode with organisationally defined risky actions held as ask rules, which removes the approval fatigue that trains developers to click through everything.
Overall, the pattern determines the dominant cost driver: DevSecOps spends on cache reads across many short sessions, while Agentic spends on regular input inside a small number of long ones.
Claude Code's context is cumulative. Every API call processes the full conversation history to date. The fixed system context is the prime candidate for prompt caching, written once per session and read cheaply on every subsequent turn.
Read deny rule on a sensitive path prevents both the selected text and the open-file notice for that file from reaching the model.
| Pattern | Use case | Tier | Cache writes | Cache reads | Regular input | Output |
|---|---|---|---|---|---|---|
| DevSecOps (Brownfield): daily token totals | ||||||
| DevSecOps | A (no MCP) | Light | 179,000 | 924,000 | 174,000 | 28,900 |
| Medium | 184,000 | 1,218,000 | 320,000 | 46,200 | ||
| Heavy | 254,000 | 1,680,000 | 467,500 | 65,200 | ||
| DevSecOps | B (with MCP) | Light | 251,000 | 1,320,000 | 226,500 | 33,400 |
| Medium | 256,000 | 1,740,000 | 393,300 | 53,100 | ||
| Heavy | 353,000 | 2,400,000 | 569,600 | 74,700 | ||
| Agentic (Greenfield): daily token totals | ||||||
| Agentic | A (no MCP) | Light | 115,000 | 945,000 | 185,000 | 30,000 |
| Medium | 62,000 | 1,008,000 | 626,000 | 40,000 | ||
| Heavy | 77,000 | 1,848,000 | 1,914,000 | 91,500 | ||
| Agentic | B (with MCP) | Light | 160,000 | 1,350,000 | 207,500 | 30,000 |
| Medium | 80,000 | 1,440,000 | 661,000 | 43,000 | ||
| Heavy | 95,000 | 2,640,000 | 1,971,500 | 96,000 | ||
Seven behaviours consume tokens outside the session shapes in section 03. Three of them bill a full-context turn while the developer is doing nothing. Budget for them at the tier level and control them at the gateway.
| Driver | Mechanism | Control |
|---|---|---|
| Compaction | /compact reads the conversation it summarises, so compacting a large context is itself a large request, and the following turn writes a fresh prefix. /clear costs nothing | Prefer /clear between unrelated tasks. Reserve /compact for continuity within one task |
/goal conditions | A separate evaluator re-checks the goal after every turn, and idle check-ins start a new turn carrying full context while background work runs | Set CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 where goals are not required |
| Stop hooks | A Stop hook blocks the turn from ending until the check passes, for up to 8 consecutive blocks | Keep the check cheap and fast. Budget up to 8 extra turns per gated task |
| Scheduled tasks | Fires on its interval and sends full context whether or not the session is active | Disable by policy unless a named use case justifies it |
| Cross-session messaging | A message from another session arrives as a new turn with full context | Set crossSessionInbound to hold |
/batch | Splits a change across 5–30 subagents, each in its own worktree, each opening a merge request | Treat as an uncapped multiplier alongside agent teams. Govern at the service budget, not by guidance |
| Auto mode classifier | A separate classifier evaluates every tool call and consumes a small number of extra tokens per call. Those calls count toward token usage on Bedrock. Auto mode also falls back to manual approvals after three consecutive blocks or twenty across a session, adding approval turns | Unavoidable while auto mode is enabled, and the safety case supports keeping it. Measure the per-tool-call overhead at the 30-day review and fold it into the tier volumes |
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1. Where a team approves them, four controls limit the exposure: Sonnet 5 for teammates, small teams, focused spawn prompts (teammates load CLAUDE.md, MCP servers and skills automatically), and shutting teammates down when their work is done. The service budget ceiling remains the primary control.
Overall, the fixed context is cheap to cache and expensive to ignore, and the six background drivers above are the difference between a modelled budget and an actual invoice.
Opus costs 2.5 times Sonnet 5 per input token, and 3.25 times once the tokenizer difference is included. Two developer tiers define access. The secd3v Claude Code Service enforces both by role-based routing, per user, per team and organisation wide.
Spec authoring opens Opus access to standard developers on amortisation grounds. A single Opus 4.8 spec session amortises across 40 to 65 Sonnet 5 execution turns. The overhead recovers within the first execution session through reduced file-exploration turns and eliminated wrong-direction corrections. The service implements a spec-write session profile that permits Opus and enforces Sonnet-only for subsequent sessions in the same sprint.
The Haiku share differs between patterns for a reason grounded in task shape. In DevSecOps a substantial fraction of daily interactions, including documentation strings, pipeline status checks, dependency lookups and boilerplate, genuinely do not require Sonnet-level intelligence. In standard Agentic sessions even simple tasks involve sustained multi-turn reasoning where Haiku creates quality drag and more wrong-direction attempts. In spec-driven Agentic, Phase 3 conformance review is pattern matching against defined criteria and Haiku handles it at a third of the cost.
tools, system and messages token count, so the 21,000 tokens of fixed context clear it comfortably in a normal session. The threshold constrains only a subagent invocation carrying a small prompt and no tool definitions, where nothing caches and no error is returned.
Sonnet 5 and Opus 4.8 improve on their predecessors on measurable grounds, and the recommendation to select them is not a preference for the newer label.
| Comparison | Where the current model is better | Where it costs more |
|---|---|---|
| Sonnet 5 against Sonnet 4.6 | Anthropic positions it as its most agentic Sonnet. The rate card is 33% lower, and after the tokenizer increase the measured spend is 13.3% lower for identical work, because input, output and both cache rates all scale by the same factor | Nothing. Sonnet 5 costs less on the rate card and less in measured spend |
| Opus 4.8 against Opus 4.6 | Minimum cache checkpoint falls from 4,096 tokens to 1,024, which widens what can be cached in short sessions. Adaptive reasoning brings effort-level control. A new system instruction can be added partway through a conversation without invalidating the system or message caches | Approximately 30% more expensive in practice at an identical rate card, because of the tokenizer |
Overall, the standard-developer split costs 18% less than the senior split on DevSecOps and 22% less on Agentic, and the difference is entirely the Opus share.
Cost per developer per month (USD), 22 working days, 5-minute TTL, Australian endpoint rates, with the tokenizer multiplier applied to Sonnet 5 and Opus 4.8. Standard = no-Opus split. Senior = limited-Opus split. Per-model pure costs shown only for building custom blends.
Each monthly figure applies the section 02 rates to the section 04 daily token volumes across 22 working days, then blends the per-model results by the section 05 splits.
Two assumptions sit behind the volumes and should be stated rather than left implicit. The opening turn of each session carries its context as a cache write rather than as regular input, which is why both the cache-read and regular-input derivations subtract the session count from the turn count. Cache writes then exceed the simple product of sessions and fixed context, because caching writes incrementally as the conversation grows, and that increment is the residual in the cache-write column. On that basis the section 04 volumes follow from the section 03 session definitions as set out below.
| Quantity | Derivation | Agreement with the published volumes |
|---|---|---|
| Turn counts | Sum of turns per session type across the tier's session mix | Exact for all six tiers |
| Output, use case A (no MCP) | Sum of turns multiplied by per-turn output for each session type | Exact for all six tiers. DevSecOps heavy: 25 × 400, plus 36 × 700, plus 30 × 1,000 = 65,200 |
| Cache reads, use case A (no MCP) | (Turns less sessions) × 21,000 | Exact for all six tiers. DevSecOps heavy: 80 read-turns × 21,000 = 1,680,000 |
| Cache reads, use case B (with GitLab MCP) | (Turns less sessions) × 30,000 | Exact for all six tiers. Agentic heavy: 88 read-turns × 30,000 = 2,640,000 |
| Regular input, use case A (no MCP) | (Turns less sessions) × per-turn regular input | Within 0.1–2.7%, the variance being a small allowance carried in the published volumes |
| Regular input and output, use case B (with GitLab MCP) | Use case A volumes plus a per-turn uplift for MCP payloads | Requires the uplift below, which the session definitions do not by themselves supply |
| Policy / Model | Split | Light / mo | Medium / mo | Heavy / mo |
|---|---|---|---|---|
| Standard developer | 3070 | $32.56 | $46.32 | $65.62 |
| Senior (limited Opus) | 256510 | $39.78 | $56.59 | $80.17 |
| Saving: Standard vs Senior | −$7.22 (18%) | −$10.27 (18%) | −$14.55 (18%) | |
| Per-model pure reference | ||||
| Haiku 4.5 | n/a | $15.36 | $21.85 | $30.95 |
| Sonnet 5 | n/a | $39.93 | $56.80 | $80.47 |
| Sonnet 4.6 | n/a | $46.08 | $65.54 | $92.86 |
| Opus 4.8 | n/a | $99.83 | $142.01 | $201.19 |
| Opus 4.6 | n/a | $76.79 | $109.24 | $154.76 |
| Policy / Model | Split | Light / mo | Medium / mo | Heavy / mo |
|---|---|---|---|---|
| Standard developer | 3070 | $43.06 | $59.14 | $83.34 |
| Senior (limited Opus) | 256510 | $52.60 | $72.26 | $101.81 |
| Saving: Standard vs Senior | −$9.55 (18%) | −$13.11 (18%) | −$18.48 (18%) | |
| Per-model pure reference | ||||
| Haiku 4.5 | n/a | $20.31 | $27.90 | $39.31 |
| Sonnet 5 | n/a | $52.81 | $72.53 | $102.20 |
| Sonnet 4.6 | n/a | $60.93 | $83.69 | $117.93 |
| Opus 4.8 | n/a | $132.01 | $181.34 | $255.51 |
| Opus 4.6 | n/a | $101.55 | $139.49 | $196.55 |
| Policy / Model | Split | Light / mo | Medium / mo | Heavy / mo |
|---|---|---|---|---|
| Standard developer | 1585 | $32.74 | $57.36 | $151.49 |
| Standard developer, spec-driven est. | 20728 | $30.39 | $44.15 | $108.76 |
| Saving: Standard vs Spec-Driven | −$2.35 (7%) | −$13.21 (23%) | −$42.73 (28%) | |
| Senior (limited Opus) | 107515 | $41.96 | $73.52 | $194.18 |
| Senior, spec-driven est. | 186715 | $33.96 | $49.34 | $121.55 |
| Saving: Standard vs Senior | −$9.23 (22%) | −$16.16 (22%) | −$42.69 (22%) | |
| Per-model pure reference | ||||
| Haiku 4.5 | n/a | $13.87 | $24.30 | $64.19 |
| Sonnet 5 | n/a | $36.07 | $63.19 | $166.90 |
| Sonnet 4.6 | n/a | $41.62 | $72.91 | $192.58 |
| Opus 4.8 | n/a | $90.17 | $157.98 | $417.25 |
| Opus 4.6 | n/a | $69.36 | $121.52 | $320.96 |
| Policy / Model | Split | Light / mo | Medium / mo | Heavy / mo |
|---|---|---|---|---|
| Standard developer | 1585 | $39.55 | $63.97 | $161.87 |
| Standard developer, spec-driven est. | 20728 | $37.19 | $50.35 | $118.42 |
| Saving: Standard vs Spec-Driven | −$2.36 (6%) | −$13.62 (21%) | −$43.45 (27%) | |
| Senior (limited Opus) | 107515 | $50.69 | $81.99 | $207.48 |
| Senior, spec-driven est. | 186715 | $41.57 | $56.28 | $132.36 |
| Saving: Standard vs Senior | −$11.14 (22%) | −$18.02 (22%) | −$45.61 (22%) | |
| Per-model pure reference | ||||
| Haiku 4.5 | n/a | $16.76 | $27.10 | $68.59 |
| Sonnet 5 | n/a | $43.57 | $70.47 | $178.33 |
| Sonnet 4.6 | n/a | $50.28 | $80.47 | $205.77 |
| Opus 4.8 | n/a | $108.93 | $176.18 | $445.83 |
| Opus 4.6 | n/a | $83.79 | $135.52 | $342.94 |
| Pattern | Use case | Light / mo | Medium / mo | Heavy / mo | Dominant cost driver at heavy |
|---|---|---|---|---|---|
| DevSecOps | A (no MCP) | $32.56 | $46.32 | $65.62 | Cache reads across 11 short sessions |
| Agentic | A (no MCP) | $32.74 | $57.36 | $151.49 | Regular input at 1,914k tok/day, 4.1 times DevSecOps |
| Spec-Driven est. | A (no MCP) | $30.39 | $44.15 | $108.76 | Regular input cut 48% by the spec; Phase 3 on Haiku |
| DevSecOps | B (with MCP) | $43.06 | $59.14 | $83.34 | MCP adds 32% at light and 27% at heavy |
| Agentic | B (with MCP) | $39.55 | $63.97 | $161.87 | Heavy agentic runs 94% above DevSecOps heavy |
| Spec-Driven est. | B (with MCP) | $37.19 | $50.35 | $118.42 | Spec-driven narrows the gap against DevSecOps to 42% at heavy |
Every figure in this document models software development: reading and writing code, reviewing changes, generating tests, auditing for security defects, and building services against a specification. The session definitions in section 03 are drawn from that work and nothing else.
Developers also use the agent for content activities that sit alongside development and are not costed above. These include long-form documentation, release notes and changelogs, merge request and release summarisation, architecture and sequence diagrams, commit message drafting, incident write-ups, runbooks, onboarding material, and open-ended explanations of unfamiliar code. Some of this work appears in the model only incidentally, through the documentation update and quick explanation examples in the DevSecOps micro session, and the rest is absent.
These sessions carry a different token shape from development work. They run short, they read a bounded amount of context, and they produce a great deal of output, which is the most expensive token category at five times the input rate. One such session per day, modelled at 6 turns with 4,000 tokens of regular input per turn and 1,500 tokens of output per turn on the standard developer split, costs approximately $5.22 per month.
| Base figure, standard policy | One session / day | Two sessions / day |
|---|---|---|
| DevSecOps light, use case A (no MCP), $32.56 | 16.0% | 32.1% |
| DevSecOps heavy, use case A (no MCP), $65.62 | 8.0% | 15.9% |
| DevSecOps heavy, use case B (with GitLab MCP), $83.34 | 6.3% | 12.5% |
| Agentic heavy, use case A (no MCP), $151.49 | 3.4% | 6.9% |
| Agentic heavy, use case B (with GitLab MCP), $161.87 | 3.2% | 6.4% |
| Pattern and tier, standard policy | Base / mo | With 10% allowance / mo |
|---|---|---|
| DevSecOps heavy, use case A (no MCP) | $65.62 | $72.18 |
| DevSecOps heavy, use case B (with GitLab MCP) | $83.34 | $91.67 |
| Agentic heavy, use case A (no MCP) | $151.49 | $166.64 |
| Agentic heavy, use case B (with GitLab MCP) | $161.87 | $178.06 |
| Spec-driven heavy, use case A (no MCP) est. | $108.76 | $119.64 |
| Spec-driven heavy, use case B (with GitLab MCP) est. | $118.42 | $130.26 |
Overall, spec-driven development delivers the largest single return on investment available in the model, cutting heavy agentic spend by 28% for the cost of one authoring session per sprint, and a 10% allowance covers the content activities these figures exclude.
Prompt caching is the largest cost lever in this model. Without it, the fixed system context bills as full-price regular input on every API call. Prompt caching is enabled by default on the Bedrock API, and the 5-minute TTL applies unless a request specifies otherwise.
Standard developer blend, 5-minute TTL, AU regional.
| Pattern | Use case | Tier | No cache / mo | 5-min TTL / mo | Saving |
|---|---|---|---|---|---|
| DevSecOps | |||||
| DevSecOps | A (no MCP) | Light | $72.93 | $32.56 | $40.37 (55%) |
| Medium | $100.20 | $46.32 | $53.88 (54%) | ||
| Heavy | $139.93 | $65.62 | $74.31 (53%) | ||
| DevSecOps | B (with MCP) | Light | $100.79 | $43.06 | $57.73 (57%) |
| Medium | $136.20 | $59.14 | $77.06 (57%) | ||
| Heavy | $189.62 | $83.34 | $106.29 (56%) | ||
| Agentic | |||||
| Agentic | A (no MCP) | Light | $79.67 | $32.74 | $46.93 (59%) |
| Medium | $108.28 | $57.36 | $50.93 (47%) | ||
| Heavy | $245.38 | $151.49 | $93.89 (38%) | ||
| Agentic | B (with MCP) | Light | $106.66 | $39.55 | $67.11 (63%) |
| Medium | $136.84 | $63.97 | $72.87 (53%) | ||
| Heavy | $296.21 | $161.87 | $134.34 (45%) | ||
AWS and Anthropic both publish the same guidance, and both frame the 5-minute cache as the default for regularly used prompts rather than as a limitation to be worked around.
A Claude Code session is a tool-use loop in which one conversational turn issues many API requests seconds apart, so the gaps that matter are genuine idle periods rather than turn boundaries. That is precisely the regular cadence the vendor guidance describes. Production telemetry agrees: cache hit rates exceed 90% on the 5-minute TTL across Claude Code and Claude Code Service usage.
| Pattern | Use case | Tier | 5-min TTL / mo | 1-hr TTL / mo | 1-hr premium |
|---|---|---|---|---|---|
| DevSecOps: the premium is largest, across eleven short sessions a day | |||||
| DevSecOps | A (no MCP) | Light | $32.56 | $39.45 | +$6.89 |
| Medium | $46.32 | $53.40 | +$7.08 | ||
| Heavy | $65.62 | $75.39 | +$9.77 | ||
| DevSecOps | B (with MCP) | Light | $43.06 | $52.71 | +$9.66 |
| Medium | $59.14 | $68.99 | +$9.85 | ||
| Heavy | $83.34 | $96.92 | +$13.58 | ||
| Agentic: close to indifferent either way | |||||
| Agentic | A (no MCP) | Light | $32.74 | $37.67 | +$4.93 |
| Medium | $57.36 | $60.01 | +$2.66 | ||
| Heavy | $151.49 | $154.79 | +$3.30 | ||
| Agentic | B (with MCP) | Light | $39.55 | $46.40 | +$6.85 |
| Medium | $63.97 | $67.39 | +$3.43 | ||
| Heavy | $161.87 | $165.94 | +$4.07 | ||
claude -p invocationsENABLE_PROMPT_CACHING_1H is a process-level environment variable, not a per-request setting. It applies to every session a developer runs and cannot be selected per session type, so this is a per-developer-profile decision.Where residual misses come from invalidation rather than expiry, a longer TTL buys nothing and costs the premium. Six documented causes apply.
| Cause | Effect | Remedy |
|---|---|---|
| Effort level changed mid-session | The resolved effort value renders into the prompt, so a change invalidates message blocks | Set effort once per session and hold it. See section 08 |
| Thinking configuration changed | Same mechanism as effort | Set once per session |
| Tools section changed | Checkpoints chain in the order tools, system, messages, so modifying tools invalidates the system and messages caches with it | Hold the managed MCP configuration stable between sessions |
| Image added | Adding an image anywhere in the prompt invalidates message blocks | Expect a full re-write after a pasted screenshot |
| 20-block lookback exceeded | Automatic prefix checking looks back approximately 20 content blocks from the checkpoint, and static content beyond that range is not found | Occurs on long parallel tool sequences. Additional checkpoints are the documented remedy, up to the four-checkpoint maximum |
| Cross-region routing under load | At times of high demand, cross-region inference optimisations may lead to increased cache writes | Prefer direct in-region routing for cache-sensitive workloads where capacity allows |
/compact also produces a full cache write on the following turn, because summarisation replaces the conversation prefix; that is a cost of compaction rather than a cache fault and section 04 carries it.
Overall, caching delivers 38–63% against an uncached baseline, and the 5-minute default captures that saving with no configuration and no premium.
After model policy and caching configuration, developer behaviour controls cost. Each factor below is validated against Anthropic's Claude Code best practices documentation.
Shift+Tab until the status bar shows plan mode on, or start the session with claude --permission-mode plan. Press Ctrl+G to open the plan in a text editor and edit it before Claude proceeds.
/goal condition or a Stop hook enforces it, and section 04 carries the token cost of both.
/compact at 80% context fill during Phase 2.
/rename before clearing, then claude --resume to return.
/compact summarises history rather than clearing it, and takes focus instructions such as /compact focus on the API changes.
/btw asks a side question whose answer never enters conversation history. The rewind menu offers summarise-from-here and summarise-up-to-here, which condense part of the conversation while leaving the rest intact.
low or medium. DevSecOps extended and agentic execution: medium. Spec-driven Phase 3: low. Phase 1 authoring and novel reasoning: high.
MAX_THINKING_TOKENS. On Sonnet 4.6 and Opus 4.6, MAX_THINKING_TOKENS applies in combination with CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING=1.
model: haiku in the subagent configuration for file scanning, documentation lookup and log analysis, and the main Sonnet 5 session receives a summary without paying Sonnet prices for the exploration.
PreToolUse hook greps for ERROR and returns only the matching lines, cutting context from tens of thousands of tokens to hundreds. For a brownfield estate with verbose CI output this is the largest untapped lever in the DevSecOps pattern.
.claudeignore on day one of any brownfield project. A single accidental glob-all read of a large repository consumes 50,000–150,000 tokens in one call, most of a day's DevSecOps budget.
claude in the integrated terminal requires the standalone CLI install on the shell PATH. Installing the graphical extension does not provide it, because the extension bundles a private copy of the CLI for its own chat panel.
claude -p (pre-commit automation and fan-out loops), claude --permission-mode plan, claude --resume, claude mcp list (the managed MCP validation check) and claude --version are all unavailable, and each is a cost-reducing behaviour elsewhere in this document.
/clear and /compact are in-session commands and do not depend on PATH, but starting the CLI does.
Overall, prompt specificity, plan mode and spec authoring carry the highest return of the twelve factors, and the remainder protect that return rather than add to it.
Opus is justified where the task requires novel reasoning under genuine ambiguity, sustained autonomous operation over many turns, or where downstream error costs are high enough that a reasoning gap materially changes outcomes. Source the current benchmark comparison from Anthropic's model cards before quoting figures in a business case.
Spec authoring introduces an Opus justification that applies to all developer tiers. A single Opus 4.8 spec session producing a tight 1,500–2,500 token specification amortises across 40–65 Sonnet 5 execution turns, and the cost recovers within the first execution session.
Overall, four scenarios justify Opus outright, two are conditional on Sonnet 5 failing first, and seven do not justify it at any tier.
| Setting | Requirement | Consequence if omitted |
|---|---|---|
| Caching | ||
ENABLE_PROMPT_CACHING_1H | Leave unset so the Bedrock 5-minute default applies. Enable per developer profile only on measured expiry, per section 07 | Enabling it estate-wide costs $3–$14 per developer per month for no measured benefit |
| Background token consumption | ||
CLAUDE_CODE_GOAL_CHECKIN_MINUTES=0 | Set unless goal-driven sessions are an approved workflow | Idle check-ins send full context on an interval |
crossSessionInbound = hold | Prevents inbound cross-session messages arriving as full-context turns | Uncontrolled turns on idle sessions |
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS | Leave unset unless approved per team, with a service budget ceiling | Approximately 7 times token usage |
| Endpoint | ||
| Standalone CLI on the managed system PATH | Per section 08 | Automation, plan mode startup and MCP validation all unavailable |
managed-mcp.json | Per the deployment guides, listing only the CCS GitLab MCP server, and held stable between sessions | Uncontrolled MCP context and tool authority, plus cache invalidation when the tools section changes |
Overall, the agent teams flag carries the largest single financial risk at roughly 7 times token usage, and the caching default carries the largest recurring one.
Claude Code sends no usage metrics to Anthropic, so attribution comes from the request path. Because every inference request passes through the CCS gateway, the gateway is the natural and complete point of measurement, and the accounting planes described in the CCS gateway high level design are where per-user consumption is recorded. The gateway also assumes its upstream role per user with session tags, which carries the same attribution into the AWS audit trail, so CloudTrail corroborates the gateway record rather than substituting for it.
Every signal below derives from request metadata recorded at the gateway, apart from the CLAUDE.md check, which is endpoint-side.
| Signal | Indicates |
|---|---|
| Caching | |
| Cache hit rate below 70% | Poor session hygiene, or one of the invalidation causes in section 07 |
| Full-prefix cache write after an idle gap | TTL expiry. The only evidence that justifies enabling the 1-hour profile for that developer |
| Full-prefix cache write with no preceding gap | Invalidation rather than expiry. Check for effort changes, tools-section changes or pasted images |
| Cache write volume rising with no change in session pattern | Routing behaviour under load where a multi-region profile is in use, per section 02 |
| Cache token counts at zero | A prompt below the minimum cacheable length for the model in use |
| Cache write records reporting a 1-hour TTL | The 1-hour setting is active on a profile that may not warrant it |
| Context discipline | |
| Regular input above 15,000 tok/turn on DevSecOps | Broad prompting, or a missing /clear |
| Regular input above 18,000 tok/turn during spec-driven Phase 2 | The spec is not being referenced and Claude is still file-exploring |
| CLAUDE.md above 3,000 tokens on a spec-driven project | Spec prose embedded in CLAUDE.md. Endpoint-side check only, because the gateway records metadata and cannot inspect the file |
| Model policy | |
| Haiku share below 15% on DevSecOps | Model discipline not applied |
Overall, customers should conduct a recurring 30-day telemetry review against the signals above, restating the estimated inputs in this model from measured data and reissuing the affected tables.