Condensed revision notes for all five CCA-F domains, structured by exam task statement — because if you can tell which task statement a question is testing, the answer is usually obvious. Global rules first, then the traps, the distractors and the mnemonics. Pass cut: a scaled 720/1000.
Lifecycle: send → inspect stop_reason → if tool_use, execute tools, append tool_result blocks as role: "user", loop → if end_turn, done.
stop_reason is the ONLY authoritative termination signal.
content[0].type == "text" — tool_use responses can begin with a text block{"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "<id>", "content": "...", "is_error": False}
]}
Parallel tool_uses → one tool_result per id, all in one user turn.
Topology: hub-and-spoke only. Coordinator at centre; subagents are spokes; ALL communication routes through the coordinator. Subagents NEVER talk to each other.
Coordinator's 7 responsibilities: decompose → select → partition scope → pass context → aggregate → iteratively refine → route for observability.
| Symptom | Root cause |
|---|---|
| Missing topics in output | Decomposition too narrow |
| Duplicated work / redundant sources | Scope partitioning sloppy |
| Shallow on every topic | Subagent prompts or tool budgets |
| Missing source attribution | Context-passing destroyed metadata |
Task tool is the only mechanism to spawn subagents. allowedTools must include "Task" or the coordinator physically cannot delegate.
AgentDefinition fields: description (used for selection — write like a function docstring), prompt, tools (subset of coordinator's), optionally model.
Parallel spawning: emit multiple Task tool_uses in a single coordinator response → executes concurrently. Sequential = separate turns = triple latency. Only safe when subagents are independent.
fork_session: independent branches from a shared expensive baseline. Use for "explore N alternatives from the same analysis."
Taskfork_sessionPROBABILISTIC ──────────────────────────► DETERMINISTIC
prompt rules → few-shot → routing → hooks / prerequisite gates
~95-99% ~98% ~98% 100%
The decision rule (memorise verbatim): If a single failure causes financial loss, security breach, regulatory violation, or safety harm → programmatic enforcement (hooks).
If a single failure causes only a bad-but-recoverable UX → prompts.
Canonical exam scenarios → all answer = hooks: refund verification, trade limits, compliance checks, manager approvals, age/permission gates, audit logging.
Multi-concern requests: decompose → investigate in parallel with shared context → synthesise unified resolution.
| Hook | Fires | Use cases |
|---|---|---|
| PreToolUse | Before tool executes | Gates, authorisation, prerequisite checks, financial limits, capturing prior state hash for audit |
| PostToolUse | After tool, before model sees result | Normalisation (timestamps, currencies, field names), redaction (PII), truncation, enrichment, result logging |
Hooks don't rank tools. Hooks gate, block, or transform. Use prompts + tool descriptions for "prefer this tool over that one."
| Pattern | When | Tell |
|---|---|---|
| Fixed Sequential Pipeline | Stable structure, each step consumes prior step's output | "Every email follows the same flow" |
| Dynamic Decomposition (typed subagents) | Varying requests, role-typed subagents (search/doc/synth) | "Same types invoked, scope varies per query" |
| Orchestrator–Worker | Uniform workers differentiated only by prompt-supplied scope | "N research workers, each handling a different sub-topic" |
| Evaluator–Optimiser | Explicit checkable quality criteria, first-pass often inadequate | "Iterate until coverage / citations / format pass" |
Discriminator (Pattern 2 vs Pattern 3): typed roles → dynamic decomposition; uniform clones → orchestrator–worker.
Evaluator–optimiser fails when: the evaluator cannot reliably distinguish good from bad. Subjective criteria → loop spins.
If a single-agent loop works, the exam's preferred answer is to not refactor to multi-agent. Multi-agent adds latency, cost, and failure surface — warranted only when single-agent genuinely cannot do the job.
| Type | Examples | Response |
|---|---|---|
| Transient | 429, 503, network blip | Retry with exponential backoff + jitter, capped |
| Permanent | 400, 401, 403, 404, malformed input | Don't retry — feed back to model or escalate |
| Logical | Tool returned empty / irrelevant results | Feed back to model with context, model changes strategy |
Blanket retry on every error type. Always classify first.
Idempotency: generate a UUID / hash of {operation + customer + amount} on the first attempt, pass it on every retry, server-side dedup. Without this, a transient retry can double-charge.
Paired-fix pattern (exam favourite): when both retry policy is broken AND duplication occurs, you need both error classification and idempotency keys. Either alone is insufficient.
Errors must reach the model when the model can do something useful. Silent harness error-swallowing (returning empty string for failed tool calls) makes the model produce confident wrong answers.
Task to coordinator's allowedTools" (for "won't delegate" bugs)stop_reason or it's wrong.allowedTools for "Task".tool_choice before prompt instructions. Environment variables before relocating files. Scoped tools before merging agents.isError: true means access failure. isError: false with empty result means success-with-no-data. Conflating them breaks all recovery logic.The core principle: tool descriptions ARE the routing mechanism. The model has nothing else when choosing between similar tools.
other_tool." ← the one most teams skip| Fix | Rank | Why |
|---|---|---|
| Expand tool descriptions | 1st — correct | Low effort, high leverage, fixes root cause |
| Few-shot examples in system prompt | 2nd | Treats symptom, token overhead on every call |
| Merge tools into one | 3rd | Loses useful semantic distinction, high refactor cost |
| Train a routing classifier | 4th — wrong | Extra model call, latency, new failure point |
The exam's preferred answer is ALWAYS "expand descriptions" as the first step for a misrouting scenario.
Code smell: a tool description containing "or" between behaviours, or "depending on" between conditions. Example:
"Processes refunds, partial refunds, store credits, or exchanges depending on policy and customer status."
This is unsplittable at runtime. The model can't route to it reliably. Split into purpose-specific tools with tight contracts:
issue_full_refund(order_id, reason)issue_partial_refund(order_id, amount, reason)issue_store_credit(customer_id, amount, reason)process_exchange(order_id, replacement_sku, reason)Renaming a generic tool does NOT fix this. Splitting does.
Even with perfect tool descriptions, misrouting can persist if the system prompt contains keyword-sensitive instructions ("Always look up the customer first") that override tool descriptions. Rule: after fixing tool descriptions, audit the system prompt for keyword conflicts. Don't add more routing rules to the prompt — that competes with the tool descriptions. Remove the conflicting keyword associations instead.
The MCP isError flag is the agent's primary recovery signal. Without structured error metadata, agents guess — and guess badly (retry things that won't succeed, escalate things that needed a moment).
| Category | Examples | Retryable? | Agent response |
|---|---|---|---|
| Transient | Network timeout, 503, rate limit | Yes (same input) | Backoff + retry, then escalate |
| Validation | Wrong format, missing field, malformed ID | Yes (after fixing input) | Correct input and retry |
| Business | Refund exceeds limit, account locked, not eligible | No | Switch workflow; surface description to user |
| Permission | 401, 403, missing scope | No (by this agent) | Escalate or request elevated creds |
{
"isError": true,
"errorCategory": "business",
"isRetryable": false,
"description": "Refund of £450 exceeds £200 self-service limit. Customer must speak with a supervisor."
}
The description for business errors should be customer-friendly — the agent will often surface it to the user.
| Case | Response | Agent action |
|---|---|---|
| Access failure (tool can't reach source) | isError: true, category transient/permission | Decide whether to retry |
| Valid empty result (tool queried successfully, found nothing) | isError: false, empty array / null result | Tell user "no results" |
Conflating these. Returning [] for both "no such customer" AND "database unreachable" causes the agent to retry endlessly and escalate with a misleading description, when the real answer was "this customer doesn't exist."
isRetryable: true = retry with the same inputisRetryable: true = retry only after correcting the inputSubagents handle local recovery. Only propagate what they cannot resolve.
Retry budget: even retryable errors have limits. A well-designed tool may include retryAfter and maxRetries hints. "Agent retried 50 times" in a setup = missing retry budget.
tool_choicetool_choice values (know each verbatim)| Need | tool_choice |
|---|---|
| Model may or may not call a tool | "auto" (default) |
| Model MUST call some tool, of its choice | "any" |
| Model MUST call this specific tool | {"type": "tool", "name": "extract_intent"} |
Scenario: "Every customer message must first extract intent and entities before routing. Currently done via system prompt instruction; ~15% of messages skip the extraction."
Correct fix: set tool_choice: {"type": "tool", "name": "extract_intent"} on the first turn. Switch to "auto" for subsequent turns. The model literally cannot output anything other than the forced tool call. 15% miss rate → 0%.
tool_choice: any" — model could pick a different toolRule: when the requirement is "the model MUST do X first", reach for tool_choice: {"type": "tool", "name": "X"}, not a prompt instruction. Prompts request; tool_choice requires.
Setup: synthesis agent frequently returns control to coordinator for one-line fact verifications. 85% are simple lookups, 15% need multi-source cross-referencing. Round-trip latency dominates.
web_search + fetch_page → defeats role scoping, tool overload, prompt injection riskCorrect fix: add a scoped, constrained verify_fact(claim, max_sources=3) tool to the synthesis agent. Tool description explicitly states that complex multi-source verifications must be routed back to the coordinator.
Pattern in one sentence: for high-frequency simple operations, give the consuming agent a constrained version of the capability; complex cases still escalate to the owner.
fetch_url(url) → load_document(document_url) — validates allowed domains, document content-types, strips scriptsrun_sql(query) → query_orders(filters) — single-table, parameterisedsend_email(...) → send_customer_notification(template_id, ...) — templated, restricted recipientsSame capability, narrower surface, reduced blast radius.
| Level | Path | Version-controlled? | Purpose |
|---|---|---|---|
| Project | .mcp.json (repo root) | Yes (committed, shared) | Team-wide integrations for this project |
| User | ~/.claude.json | No (personal) | Personal/experimental servers across all projects |
Tool discovery happens at connection time — all tools from all configured servers are made available simultaneously. Too many configured servers re-creates the tool-overload problem at a different layer.
${VAR} expansion (most-tested 2.4 pattern)Wrong:
{ "env": { "JIRA_API_TOKEN": "ATATT3xFf_actual_token_here" } }
Right:
{ "env": { "JIRA_API_TOKEN": "${JIRA_API_TOKEN}" } }
Rule: "shared config, personal credentials". The
.mcp.jsonstructure stays committed and shared. Secrets live in each developer's local environment.
Don't fix a credential leak by relocating the file. A team integration belongs in .mcp.json so new teammates have it. The fix is parameterisation, not relocation. Also: rotate any leaked token — git history is forever.
Example: jira://projects/INGEST/summary, db://schemas/orders, github://repos/acme/api/issues/labels.
Exam framing: "Agent makes many exploratory tool calls at conversation start to discover available data" → expose resources that catalogue what's available. The agent reads the catalogue once, then queries precisely.
Use a community MCP server when:
Build custom only when:
Exam preference: "Configuration over construction" and "fork over fresh build". When a community server is 90% of what you need (e.g., just want different return fields), post-process its responses or fork it — don't write from scratch.
When MCP tools overlap with Claude Code built-ins (Grep, Glob, Read), the agent may reach for the built-ins because their descriptions are tighter. Fix: enhance the MCP tool's description with examples and explicit boundary clauses (same principle as 2.1). E.g., "Use this for semantic queries like 'where do we handle payment retries'. Prefer Grep for exact-string searches like 'TODO: fix'."
| Tool | Searches | Returns | Mnemonic |
|---|---|---|---|
| Grep | File contents | Matching lines + filenames | "What does the code say?" |
| Glob | File paths | File paths matching pattern | "What files exist?" |
Grep use cases: function callers, error messages, import statements, TODO comments, identifier occurrences.
Glob use cases: all test files (**/*.test.tsx), all configs (**/*.config.{js,ts,json}), files by extension.
**/*.tsx, NOT Grep import ReactformatDate" → Grep formatDate(, NOT Glob (can't match function calls with a path pattern)| Tool | Cost | When |
|---|---|---|
| Read | Tokens proportional to file size | When you need to see contents |
| Edit | Low (diff-only) | Targeted modification via unique text anchor |
| Write | Full file in context | Complete rewrites or Edit fallback |
old_string anchor with surrounding context until unique — the FIRST movereplace_all: trueDistractors to reject: "use sed" (wrong tool category), "delete and recreate" (destructive).
Loading ten 500-line files burns tens of thousands of tokens before any reasoning. Context budget destroyed.
Correct pattern:
Wrapper-module tracing: for codebases with re-exports (index.ts does export { formatDate } from './date'), Grep finds both definitions and re-exports. Identify which references are definitions vs. re-exports vs. call sites, then trace the chain.
Task: "Find all callers of legacyAuth, then find the test files for those callers."
Correct sequence:
legacyAuth( → returns caller files (userService.ts, paymentService.ts, …)**/{userService,paymentService}.test.{ts,tsx} → returns test filesRule: Grep first when you're discovering content; Glob second to narrow paths around the results. Mirror form: known file set → check contents goes Glob → Grep. The order depends on what is the seed vs. the filter.
tool_choice: {"type": "tool", "name": "X"})PreToolUse hook to enforce…" (Claude Code harness feature, NOT Anthropic API agent design — wrong layer).mcp.json to ~/.claude.json to fix the credential leak…" (use ${VAR} expansion in the shared file instead).mcp.json to .gitignore…" (defeats team sharing)web_search…" (defeats role scoping)isError: false with an empty result for the valid-no-data case"tool_choice: {"type": "tool", "name": "X"} to force the mandatory first step"verify_fact tool to the consuming agent with explicit fallback to the coordinator for complex cases"fetch_url with load_document that validates the URL surface"${VAR} expansion in .mcp.json; rotate the leaked token"isError: true: the tool couldn't look. isError: false + empty: the tool looked and found nothing.tool_choice: {"type": "tool", "name": "X"} — never a prompt instruction.load_document not fetch_url)..mcp.json: ${VAR} expansion + rotate the leaked token. Don't relocate the file.PreToolUse etc.) are CLI-harness features. API-level agent design uses tool_choice, tools, and structured tool results. Don't reach for a hook on an API-design question.Many wrong answers come from reaching for the right idea at the wrong layer:
| Layer | Tools available |
|---|---|
| Anthropic API agent design | tool_choice, tools, system prompt, tool result structure, isError |
| Claude Code CLI harness | PreToolUse, PostToolUse, Stop hooks, settings.json |
| MCP server config | .mcp.json, ~/.claude.json, ${VAR} expansion, resources |
| Tool implementation | Error categories, retry hints, input validation, scoped surface |
When you feel the pull toward a fix, ask which layer the question is testing. The exam is at the API + MCP + tool-design layers. Hooks are usually the wrong answer in Domain 2.
~/.claude/ is personal, never shared. .claude/ (project) is shared via git. Same filename, opposite scope. Memorise the paths cold; reasoning won't save you..claude/rules/ with paths: frontmatter. On-demand workflows live in skills/commands. Don't bloat one mechanism with another mechanism's job.| Level | Path | Scope | Version-controlled? |
|---|---|---|---|
| User | ~/.claude/CLAUDE.md | Only you, on your machine | No |
| Project | .claude/CLAUDE.md or root CLAUDE.md | Everyone on the repo | Yes |
| Directory | <subdir>/CLAUDE.md | Only when working in that subdir | Yes |
Directory CLAUDE.md scopes downward and local, never upward. A src/backend/CLAUDE.md does NOT affect work in src/.
Setup: "Developer A's Claude Code follows team conventions. Developer B, on the same repo, same branch, gets inconsistent behaviour."
Root cause 9/10 times: Developer A wrote the conventions to ~/.claude/CLAUDE.md (user-level, lives in their home directory). Git never carried it. Developer B's clone has nothing.
Diagnostic instinct: Divergent behaviour between teammates on the same repo → suspect user-level vs project-level config first.
Trap distractors: "Run /memory reload" — can't refresh a file that doesn't exist on the machine. "Restart Claude Code" — same problem. "Move into directory-level CLAUDE.md" — wrong tool for cross-cutting standards.
@import syntax inside CLAUDE.md — reference external files, import per-package standards.claude/rules/ directory — topic-specific files (testing.md, api-conventions.md, deployment.md) as an alternative to one massive CLAUDE.md/memory command — the diagnostic toolLists which memory files are currently loaded. Use to confirm the diagnosis, not to fix the bug. If a file you expected isn't there, you've found your bug — but the fix is to put the file in the right location, not to run /memory harder.
| Type | Location | Shared? |
|---|---|---|
| Project slash commands | .claude/commands/ | Yes — version-controlled |
| Personal slash commands | ~/.claude/commands/ | No |
| Project skills | .claude/skills/<name>/SKILL.md | Yes |
| Personal skills | ~/.claude/skills/<name>/SKILL.md | No |
Each skill is its own directory containing a SKILL.md file. The directory name becomes the skill name.
---
name: brainstorm
description: Explore design alternatives for a feature
context: fork
allowed-tools: [Read, Grep]
argument-hint: <feature-name>
---
| Option | Effect | Exam trigger phrase |
|---|---|---|
context: fork | Runs in isolated sub-agent context; verbose output stays in fork; only summary returns to main conversation | "produces verbose output", "clutters main conversation", "keep main context clean" |
allowed-tools | Restricts which tools the skill can call; prevents destructive actions | "analyse only, must not modify", "restrict to read-only" |
argument-hint | Prompts the developer for required parameters when invoked without arguments | "prompt the user for the parameter" |
| Skills | CLAUDE.md | |
|---|---|---|
| When loaded | On-demand, when invoked | Always loaded |
| Purpose | Task-specific workflows | Universal standards |
| Example | /review, /migrate-component | "Use British English in comments" |
context: forkScenario: Team has a /review skill; one developer wants a stricter variant for personal use.
~/.claude/skills/review-strict/SKILL.md — different name, personal location, no impact on others..claude/ vs ~/.claude/)context: fork or notA skill can be project-scoped AND forked. Personal AND unforked. These are orthogonal.
Mechanism: files in .claude/rules/ with YAML frontmatter declaring a glob:
---
paths: ["**/*.test.tsx", "**/*.test.ts"]
---
All test files must use describe/it blocks, not test().
Mock external services using shared fixtures in tests/fixtures/.
Loads only when Claude Code is about to work on a file matching the glob. Not loaded for non-matching files.
| Directory CLAUDE.md | Path-specific rules | |
|---|---|---|
| Scope | One directory only | Anywhere the glob matches |
| Spread across codebase | Need one CLAUDE.md per directory | One rule file, globs everywhere |
| Token cost | Loaded whenever working in that directory | Loaded only on matching files |
Glob patterns match files scattered anywhere. **/*.test.tsx catches every test file regardless of directory. A directory-level CLAUDE.md cannot do this — you'd need N copies that drift.
| Need | Tool |
|---|---|
| Universal standards (always relevant) | Project-level CLAUDE.md |
| Standards for one specific subdirectory | Directory-level CLAUDE.md |
| Standards for files spread across the codebase by pattern | Path-specific rules |
| Task-specific workflow, invoked on demand | Skill |
The signature phrase to recognise: "pattern of files spread across the codebase" → path-specific rules.
Examples that all fit the pattern: **/*.test.*, **/migrations/**/*, **/*.tf, **/api/handlers/**/*.ts.
A slash command for the convention. Slash commands require developers to remember to invoke them. Standards must be ambient, not opt-in.
This is a judgement task statement. The exam hands you a task description; you classify.
The dividing line: Are there decisions still to make, or just a known thing to do? Decisions → plan mode. Known thing → direct execution.
| Property | Effect |
|---|---|
| Isolates verbose discovery output | Search results, file listings stay in subagent |
| Returns summaries to main conversation | Preserves main context window |
| Used during multi-phase tasks | Prevents context window exhaustion |
Exam phrasing: "main conversation filling up with grep output and file listings" → Explore subagent.
Same principle as context: fork for skills — keep noise out of the main conversation. Different mechanism, same goal.
You don't have to pick one mode for the whole task. Common workflow:
If a question describes "exploring, then implementing the chosen approach" — that's the hybrid, not a single-mode choice.
| Task | Mode |
|---|---|
| Restructure monolith into microservices (boundaries undecided) | Plan mode |
Fix null pointer in UserService.getById(), stack trace clear | Direct execution |
Migrate from winston to pino across 30 files (API differs) | Plan mode, then direct execution |
| Add a date-format validation conditional to one function | Direct execution |
| React v17 → v18 across 200+ components, concurrent features undecided | Plan mode (then direct execution per file) |
Calling a multi-file migration "mechanical" when the libraries have different APIs. Scale + interface differences = plan mode.
Technique hierarchy — what to reach for when prose isn't working:
When prose descriptions get interpreted inconsistently, show 2–3 concrete examples:
Input: user_name → Output: userName
Input: api_key → Output: apiKey
Input: http_url → Output: httpUrl
The model generalises from examples more reliably than from descriptions. Three examples eliminate ambiguity that prose admits.
Exam phrasing: "Claude Code interprets the instruction differently each iteration" → concrete input/output examples.
Write tests first. Share failures. Claude iterates against a machine-checkable specification rather than your prose. Strongest when behaviour is well-defined but implementation is open.
Have Claude ask questions first before implementing. Surfaces considerations you would miss in unfamiliar domains.
Exam phrasing: "developer working in unfamiliar domain, wants to avoid missing edge cases" → interview pattern.
| Issue type | How to feed back |
|---|---|
| Fixes interact (changing one affects others) | Single message — Claude sees all together, makes consistent decisions |
| Issues are independent (orthogonal) | Sequential — fix one, verify, move on |
THE headline rule: when prose fails repeatedly, switch modalities, don't iterate on the prose. Rewording prose tends to produce different misinterpretations, not fewer. Examples constrain the output shape directly.
This is the most memorisation-heavy task statement in the domain. Drill the flags.
-p flag (print mode) — single most-tested CI detailclaude -p "<prompt>" runs in non-interactive mode-p, the CI job hangs waiting for interactive input-p| Flag | Effect |
|---|---|
--output-format json | Produces JSON instead of prose |
--json-schema <schema> | Constrains JSON to a specific shape for reliable downstream parsing |
Exam phrasing: "automated system needs to post findings as inline PR comments" → both flags together.
The same Claude session that generated code is LESS effective at reviewing its own changes.
It retains the reasoning context that produced the code → less likely to question its own decisions (motivated reasoning).
The exam-correct pattern: use an independent review instance — a fresh Claude Code session with no prior context for the review step.
The "efficient" answer feels like "reuse the session, it already knows the code." Wrong. Review needs fresh eyes.
When a PR gets new commits and you re-run the review:
Without this: developers get duplicate comments on every push → lose trust → start ignoring the bot.
Exam framing: "developers ignoring the bot's comments" / "comment fatigue" / "duplicate findings on every push" → incremental review context, NOT "remove the bot".
CI-invoked Claude Code reads the same .claude/CLAUDE.md. For test generation especially, document:
Without this: CI-generated tests are low-value boilerplate (testing getters, trivial assertions). With it: tests target meaningful behaviour.
Exam phrasing: "CI generates low-quality tests" → strengthen .claude/CLAUDE.md with testing standards.
When using CI for test generation, provide existing test files in context alongside the source file. Without them, Claude proposes scenarios already covered — duplicating tests already in the suite.
| Problem | Fix |
|---|---|
| Low-quality tests (boilerplate, getters) | Document valuable-test criteria and fixtures in .claude/CLAUDE.md |
| Duplicate tests (scenarios already covered) | Provide existing test files in the prompt context |
These look similar but have different solutions. The exam may present both options — pick based on which problem is described.
These are the five patterns that most commonly produce wrong answers on CI/CD questions. Memorise the symptom → fix mapping:
| Symptom | Root cause | Fix |
|---|---|---|
| CI job hangs indefinitely | Claude Code waiting for interactive input | Add -p flag |
| PR comments need to be posted automatically as inline comments | Output is narrative prose, not structured data | --output-format json --json-schema |
| Duplicate comments on every new commit | No context about what was already reviewed | Include prior findings + instruct "report only new or unaddressed issues" |
| Review misses issues in code that wasn't changed in the latest commit | Running only the incremental diff | Still run full review; prior-findings context handles deduplication |
| CI generates duplicate tests | Doesn't know what's already covered | Provide existing test files in context |
What blocks CI most on the exam: the -p flag. If you see "hangs" or "waiting for input" — that is the answer, always.
Same principle, two domains. On a CI/CD question it maps to -p + separate session. On a prompt-engineering question it maps to multi-instance architecture.
/memory reload to pick up the project's instructions" (when the file doesn't exist on the new dev's machine — it's user-level config that was never shared)~/.claude/CLAUDE.md to make it load on demand" (wrong direction — that removes it from sharing)/migrate slash command everyone invokes before writing a migration" (standards must be ambient, not opt-in — use path-specific rules)** globs).claude/rules/ with paths: frontmatter)-p fixes the root cause)--output-format text so output isn't buffered" (the hang isn't an output problem — -p)~/.claude/CLAUDE.md and were never version-controlled"~/.claude/CLAUDE.md to .claude/CLAUDE.md (or root CLAUDE.md) and commit".claude/rules/ with paths: frontmatter"~/.claude/skills/ with a different name to avoid affecting teammates"context: fork to the skill so verbose output stays out of the main conversation"allowed-tools: [Read, Grep] to prevent destructive actions"paths: ["**/*.test.*"] to apply test conventions across the codebase"paths: ["**/migrations/**/*"] for migration conventions everywhere"-p flag to run non-interactively"--output-format json with --json-schema for machine-parseable findings".claude/CLAUDE.md for CI to read"~/.claude/ = personal. .claude/ = shared./memory is for diagnosis, not for fix. It tells you what's loaded; it can't conjure a missing file..claude/rules/ with paths: frontmatter.~/.claude/skills/. Never edit the team skill.context: fork for noisy skills. Explore subagent for noisy investigations. Same goal, different mechanism.-p flag. Always.--output-format json --json-schema..claude/CLAUDE.md with test-value criteria and fixtures.When a question describes a need, ask in order:
.claude/rules/ with paths: glob.claude/skills/ (or command in .claude/commands/)~/.claude/ instead of .claude/.claude/CLAUDE.md (still the project-level file — CI reads the same hierarchy)If the answer to all of 1–4 is "no," you probably don't need configuration — you need a one-off prompt.
Many wrong answers come from reaching for the right idea at the wrong layer:
| Layer | What lives here |
|---|---|
User-level config (~/.claude/) | Personal CLAUDE.md, personal commands, personal skills |
Project-level config (.claude/) | Shared CLAUDE.md, shared commands, shared skills, path-rules, .mcp.json |
| Directory-level config | Subdirectory CLAUDE.md (downward-scoped) |
| Runtime flags | -p, --output-format json, --json-schema |
| Skill frontmatter | context: fork, allowed-tools, argument-hint |
| Workflow mode | Plan mode, direct execution, Explore subagent |
When you feel the pull toward a fix, ask which layer the question is testing. A CI hang is a runtime-flag question (-p), not a CLAUDE.md question. A divergent-behaviour bug is a path-prefix question (user vs project), not a /memory question. Don't cross layers.
The core principle: Specific categorical criteria obliterate vague confidence-based instructions.
| Wrong (vague) | Right (categorical) |
|---|---|
| "Be conservative." | "Flag comments only when claimed behaviour contradicts actual code behaviour." |
| "Only report high-confidence findings." | "Report bugs and security vulnerabilities. Skip minor style preferences and local patterns." |
Structure to memorise: What to flag. What to report. What to skip. Three categories, no adjectives about confidence.
When one category has a high false-positive rate, developers stop trusting all categories, including the accurate ones.
The counterintuitive fix: temporarily disable the noisy category entirely while improving its prompt. Don't tighten in place — pull it from output, restore trust in what remains, then re-enable.
Prose definitions of severity drift across runs. Anchor severity levels with actual code examples for each level.
Critical: SQL injection from unvalidated input
Example: db.query("SELECT * FROM users WHERE id=" + req.params.id)
Medium: Missing error handling on non-critical path
Example: fs.readFile(path, cb) where cb ignores err parameter
A "comprehensive rubric" of five severity levels described in prose. Looks thorough — fails because it lacks code anchors.
The headline rule: Few-shot examples are the most effective technique for consistency. Not more instructions. Not confidence thresholds. Not temperature changes. Examples.
| Property | Requirement |
|---|---|
| Count | 2–4 targeted examples for ambiguous scenarios (not 10, not 1) |
| Content | Each example shows the reasoning for why one action was chosen over plausible alternatives |
| Purpose | Generalisation to novel patterns, not pattern-matching pre-specified cases |
Documents vary in structure: inline citations vs bibliographies, narrative prose vs structured tables. Few-shot examples covering varied structures dramatically reduce hallucination because the model learns the shape of valid extractions across formats, rather than inventing data to fit the schema.
Signature phrase to recognise: "inconsistent output across runs despite detailed instructions" → few-shot examples.
| Approach | Guarantees |
|---|---|
| Prompt-based JSON ("return JSON like...") | Nothing. Model can produce malformed JSON. |
| tool_use with JSON schema | Eliminates syntax errors entirely. |
That is the only thing tool_use guarantees. Syntactic validity. Nothing else.
A strict schema with all-required fields is a fabrication factory. Required doesn't guarantee correctness — it guarantees something will be present, even if invented.
tool_choice — three modes the exam tests| Value | Behaviour | When to use |
|---|---|---|
"auto" (default) | Model may return text instead of calling a tool | When the model legitimately might have nothing to extract |
"any" | Must call a tool, model picks which | Guaranteed structured output, unknown document type, multiple tools available |
{"type": "tool", "name": "..."} | Must call this specific tool | Force a mandatory first step or single-tool workflow |
Mental shortcut: named-tool forces a specific tool; "any" forces some tool from the menu.
Named-tool offered as a distractor when document type varies. Named-tool routes heterogeneous documents through one tool that doesn't fit most of them.
| Technique | Failure mode it addresses |
|---|---|
| Optional / nullable fields | Source legitimately lacks the information → prevents fabrication |
"unclear" enum value | Ambiguous cases → lets the model express uncertainty within the schema |
"other" + freeform detail string | Categories outside the enum → captures reality instead of forcing wrong-bucket |
| Format normalisation rules in the prompt | Schemas validate types; prompts specify formats (e.g., ISO-8601 dates) |
The model uses the error to self-correct. Dramatically more effective than naive retry, because the model now has signal about what specifically went wrong.
| EFFECTIVE for | INEFFECTIVE for |
|---|---|
| Format mismatches (date format wrong) Structural output errors (object vs array) Misplaced values (value in wrong field) | Information genuinely absent from the source |
Memorise as a sentence: Retry handles output errors. Schema design handles source gaps.
A scenario where 15% of documents lack a field and the option says "add validation-retry to recover the missing data." Wrong — the data isn't missing from the extraction, it's missing from the source. Retry will produce fabrications, not corrections. The correct fix is nullable field.
detected_pattern fields — systematic improvementAdd a detected_pattern field to each structured finding (e.g., "unvalidated_user_input_in_sql_concatenation").
| Capability unlocked | Why it matters |
|---|---|
| Analyse dismissal patterns | If pattern X is dismissed 80% of the time, you have data to improve its prompt |
| Convert black-box tool to measurable system | Improvement based on evidence, not anecdote |
Exam framing: "how do you improve a code review bot's accuracy over time" → systematic data path via detected_pattern, not "iterate on the prompt".
stated_total AND calculated_total (sum the line items yourself). Disagreement flags OCR errors, fraud, or model arithmetic mistakes — three things schemas can't catch.conflict_detected boolean. When source data is internally inconsistent (header date ≠ footer date), surface the conflict rather than silently picking one.Shared principle: extract redundantly and let inconsistencies bubble up.
CI/CD note: Scenario 5 (Claude Code for CI) is a primary domain for Domain 4. The batch API and multi-pass review questions in this domain appear in CI/CD framing. The rules below apply directly.
custom_id for correlating request/response pairsThe fourth point is the sleeper trap. Agentic workflows (model calls tool, sees result, calls another tool) cannot run on batch.
| Workflow type | API choice |
|---|---|
| Blocking (someone or something is waiting) | Synchronous |
| Latency-tolerant (no human waiting) | Batch |
Blocking examples: pre-merge CI checks, IDE inline suggestions, customer-facing chat.
Latency-tolerant examples: overnight test generation, weekly compliance audits, monthly competitive intelligence.
The correct answer is never "migrate everything to batch." It is always the mixed strategy: batch the latency-tolerant workloads, keep blocking workloads synchronous. 50% saving on API cost is dwarfed by the cost of developers blocked waiting for a 6-hour batch return.
Special case — multi-turn tool calls: even if the workflow looks latency-tolerant, multi-turn agentic loops cannot run on batch. Must be synchronous.
| Pattern | What it does |
|---|---|
Identify failures by custom_id, resubmit only failures with modifications | Don't re-batch the whole 100,000; chunk oversized docs and resubmit just the failures |
| Refine prompts on a sample set BEFORE batch processing | Iteration loops in batch are ≤24h each; validate synchronously first |
The self-review limitation — foundational principle: A model reviewing its own output in the same session retains the reasoning context that produced the output. It is biased toward its prior conclusions and less likely to question its own decisions.
The fix: spawn an independent instance — fresh session, no prior context — for the review step. The cold instance catches subtle issues the original misses.
"Ask the same model to review its work before returning." Self-review within session. Wrong. Use a fresh instance.
This is the same principle as the CI/CD session-isolation rule in Domain 3.6 — different scenario, same architecture.
| Pass | Purpose |
|---|---|
| Per-file local analysis (one pass per file, focused context) | Consistent depth per file |
| Separate cross-file integration pass | Catches data flow issues, contract mismatches across files |
"Use a larger context window so the whole codebase fits in one pass." Context size doesn't fix attention dilution. Bigger window ≠ more attention per file.
| Step | What to do |
|---|---|
| Model self-reports confidence per finding | Numeric or categorical |
| Route low-confidence findings to human review | Auto-apply high-confidence findings |
| Calibrate thresholds using a labelled validation set | Don't pick by intuition |
"Set confidence threshold to 0.9 to ensure quality" without mention of calibration. 0.9 in one system is permissive; in another it's prohibitive. Always calibrate empirically.
tool_choice: 'auto' when structured output is required" (lets the model escape into prose){type: 'tool', name: 'extract_invoice'} for documents of unknown type" (named-tool ≠ heterogeneous routing)custom_id)'unclear' enum value for ambiguous cases" (express uncertainty within the schema)'other' enum value plus a freeform detail string" (extensible categorisation)tool_choice to 'any' so the model must call one of the available tools" (guaranteed structured output across document types){type: 'tool', name: '...'} to mandate a first step" (single-tool or mandatory-step workflows)detected_pattern field to enable systematic analysis of dismissal patterns" (improvement over time)calculated_total alongside stated_total and flag discrepancies" (self-correction)conflict_detected boolean for internally inconsistent source data" (surface ambiguity)custom_id and resubmit only those with modifications" (efficient batch recovery)"auto" for optional. "any" for must-call-something. Named for must-call-this."unclear" enums, "other" categories — the anti-fabrication toolkit.detected_pattern turns a black-box bot into a measurable system.custom_id correlates.custom_id, not the whole batch.When a question describes a failure, ask in order:
"unclear" enum value or "other" + freeform detailconflict_detected boolean, surface ambiguitydetected_pattern field + analysis of dismissal patternsMany wrong answers come from reaching for the right idea at the wrong layer:
| Layer | What lives here |
|---|---|
| Prompt content | Criteria, severity definitions with code examples, format normalisation rules |
| Few-shot examples | Judgement consistency, varied-structure generalisation, hallucination reduction |
| Schema design | Anti-fabrication (nullable, unclear, other), required vs optional |
| tool_choice | Whether and which tool gets called |
| Validation-retry | Recovering output errors |
| Pipeline architecture | Multi-pass review, calculated vs stated, conflict booleans |
| API selection | Synchronous vs batch based on blocking vs latency-tolerant |
| Multi-instance | Fresh sessions for independent review |
When you feel the pull toward a fix, ask which layer the question is testing. A fabrication problem is a schema-design question (nullable), not a prompt-content question ("don't fabricate"). A batch-failure question is a custom_id question, not a retry question. A self-review limitation is a multi-instance question, not a prompt-content question. Don't cross layers.
CI/CD questions span both Domain 3 and Domain 4. When a question is framed as "Claude Code for CI/CD," the answer might live in either domain. Use this map to route quickly:
| Problem type | Domain | Fix |
|---|---|---|
| Pipeline hangs / waits for input | D3 | -p flag |
| Need machine-parseable PR comments | D3 | --output-format json --json-schema |
| Duplicate comments after new commits | D3 | Prior findings in context + "report only new issues" |
| CI generates duplicate tests | D3 | Provide existing test files in context |
| Review generates low-quality tests | D3 | Strengthen .claude/CLAUDE.md with test criteria |
| Vague instructions don't reduce false positives | D4 | Categorical criteria: "flag X, skip Y, report Z" |
| Instructions produce inconsistent output | D4 | 2–4 few-shot examples with reasoning |
| Pre-merge check vs overnight report — which uses batch? | D4 | Pre-merge = sync (blocking). Overnight = batch (latency-tolerant) |
| Large PR review gives inconsistent depth per file | D4 | Per-file local passes + separate cross-file integration pass |
| Generated code reviewed by the generator session | D4 | Independent instance (fresh session, no prior context) |
The one rule that appears in both domains in different forms: An agent/session that produced output should not review that same output. Use an independent instance. In Domain 3 it's framed as "CI code review session." In Domain 4 it's framed as "self-review limitation." Same answer.
The headline rule: Summarisation is lossy compression that disproportionately destroys high-value tokens.
| Token type | Example loss |
|---|---|
| Numerical values | "$247.83 refund" → "a refund" |
| Dates | "March 3rd" → "recently" |
| Order/case IDs | "#8891" → omitted entirely |
| Percentages / quantities | "32% increase" → "an increase" |
| Customer-stated expectations | "by Friday" → vanished |
Extract transactional facts into a structured block. Inject verbatim into every prompt. Summarisation never touches it. Conversation history above and below can be compressed; the facts block is sacred.
<case_facts>
- Customer ID: 4471
- Order: #8891, placed 2026-03-03
- Refund requested: $247.83
- Customer deadline: by Friday 2026-05-22
- Channel preference: email
</case_facts>
Transformer attention is empirically biased toward the beginning and end of long inputs. Information buried in the middle is measurably less likely to be retrieved correctly. This is not fixable with a better model — it is structural.
| Mitigation | Why |
|---|---|
| Place key findings summaries at the beginning | Beginning gets reliable attention |
Use explicit section headers (## Findings, ## Sources) | Structural anchors aid navigation |
| Surface relevant slice of long tool results at the top, full payload below | Critical content escapes the middle |
A 40-field order lookup when 5 fields are needed = 35 fields of noise. Trim at the integration layer, before the result hits the context window. Untrimmed results accumulate across turns — turn 8 still carries the noise from turns 1–7.
Define a per-tool response shape with only the fields downstream reasoning needs. Treat tool responses like API contracts, not data dumps.
Claude API calls are stateless. Every request must include the complete conversation history. Drop earlier messages → break coherence. Use prompt caching on the stable prefix (system prompt + tools + early conversation) to make this economical.
In multi-agent systems, the cheapest place to fix downstream context bloat is upstream. Modify subagents to return:
Not verbose prose, not reasoning chains, not narrative. Reasoning happens inside the subagent's context; only distilled conclusions cross the boundary.
The headline rule: Three triggers warrant escalation. Two superficially plausible triggers do not.
| Trigger | Required action |
|---|---|
| Customer explicitly requests a human | Escalate immediately. Do not attempt resolution first. |
| Policy exception or gap (request outside documented policy) | Escalate — the agent cannot invent policy |
| Inability to make meaningful progress | Escalate after legitimate avenues tried |
| Bad trigger | Why it fails |
|---|---|
| Sentiment-based (customer is frustrated) | Frustration does not correlate with case complexity |
| Self-reported model confidence | Model is often incorrectly confident on hard cases and uncertain on easy ones |
| Situation | Correct action |
|---|---|
| Straightforward issue + frustrated customer (no human request) | Acknowledge frustration, offer resolution, do not escalate yet |
| Customer reiterates preference for a human after you offer help | Now escalate |
| Customer explicitly says "I want a human" up front | Escalate immediately — even if the issue looks trivially resolvable |
Memorise the distinction: Explicit human request = immediate escalation. Frustration alone = offer help first.
Search returns multiple matches for "John Smith." The correct move is always:
Ask for an additional identifier — email, phone, order number, postcode.
The cost of asking is one extra turn. The cost of acting on the wrong account is enormous (wrong refund, wrong PII exposure, wrong cancellation). Never resolve ambiguity with heuristics when the customer can disambiguate it in one turn.
The headline rule: How a failure is reported matters as much as that it happened. Silent suppression and workflow termination are both anti-patterns.
| Field | Purpose |
|---|---|
| Failure type | transient / validation / business / permission |
| What was attempted | Specific query, parameters, endpoint, request shape |
| Partial results | Anything gathered before the failure |
| Potential alternatives | Suggested retry, fallback path, narrower query, cached endpoint |
Together these give the model enough context to recover. Unstructured errors force the model to give up or hallucinate.
{"results": [], "status": "ok"} on tool failure. Orchestrator believes no data exists. Zero recovery path — this is the worst answer in any 5.3 question.Two superficially identical responses, opposite correct behaviours.
| Type | Meaning | Correct response |
|---|---|---|
| Access failure | Tool could not reach the data source (network, auth, timeout) | Consider retry or alternative path |
| Valid empty result | Tool reached the source, ran the query, found no matches | No retry. The empty result is the answer. |
The structured error type distinguishes these. Treat all empty results the same and you either retry valid-empties forever or accept access failures as truth. Both wrong.
"Section on geothermal energy is limited due to unavailable journal access (3 of 7 expected sources unreachable)."
Better than silently omitting the section. The consumer of the synthesis can now decide whether to act on partial data or wait for the gap to be filled. Silent omission destroys this signal.
The headline rule: Extended agentic sessions degrade. Verbose discovery output dilutes early findings; the model reverts to generic "typical patterns" instead of session-specific facts.
Both symptoms have the same root: high-signal early findings get diluted by high-volume later exploration.
| Strategy | Best for |
|---|---|
| Scratchpad files — write key findings to a file, reference for subsequent questions | Findings you'll need later; survives compaction and crashes |
| Subagent delegation — spawn a subagent for a specific investigation; coordinator receives distilled result only | Parallel deep dives; keeps coordinator context focused on coordination |
| Summary injection — summarise findings from one phase before spawning the next phase's subagents | Phase boundaries in multi-phase workflows |
/compact — reduce context usage when verbose output fills the window | Recovery valve when nearing the limit; trade-off: loses fidelity |
Each agent exports structured state to a known file location (a manifest) at meaningful checkpoints. On resume, the coordinator loads the manifest and injects it into the resumed agent's prompts.
The manifest is the source of truth, not the conversation history. Robust to:
Without a manifest, a crash means starting over. With a manifest, a crash means resuming from the last checkpoint.
The headline rule: Aggregate metrics hide stratum-level failures. Calibration requires ground truth, not intuition.
"Our extraction system achieves 97% accuracy. We can automate."
The trap: 97% overall accuracy can hide a 40% error rate on a specific document type that is rare in the validation set. If that document type is the high-value commercial contracts, your automation is dangerous despite the headline number.
Discipline: Always validate accuracy by document type and by field segment before automating. A single aggregate number is never sufficient evidence to remove humans from the loop.
Even after calibration and deployment, continue sampling high-confidence extractions for ongoing verification.
| Reason | Why it matters |
|---|---|
| Novel error patterns emerge over time | New document templates, edge cases, input distribution drift |
| Failures slip through silently if not sampled | By definition, they don't trigger existing review thresholds |
| Stratified across types, fields, confidence bands | Ensures coverage of all segments, not just the loud failures |
The discipline: a small, ongoing audit of "the cases we think are fine" is what catches the failures you didn't anticipate.
| Step | Detail |
|---|---|
| 1. Per-field confidence (not per-document) | Fields within a document vary wildly in difficulty |
| 2. Calibrate thresholds on a labelled validation set | Ground truth. Raw model confidence values are not directly meaningful — "0.85" might mean 70% correct or 95% correct depending on field and document type |
| 3. Route low-confidence fields to human review | The rest auto-process |
| 4. Prioritise scarce reviewer capacity on highest-uncertainty items | Reviewer time is the limiting resource |
Connection to 5.2: This is what makes confidence-based decisions reliable. Naive self-reported confidence is unreliable. Calibrated, field-level, validation-set-grounded confidence is reliable. The exam tests this distinction.
The headline rule: Attribution must travel with every claim as structured data. Conflicts are surfaced, not resolved. Dates distinguish trends from contradictions.
| Field | Purpose |
|---|---|
| Claim | The assertion itself |
| Source URL | Direct link |
| Document name | Human-readable identifier |
| Relevant excerpt | The actual passage supporting the claim |
| Publication date | Temporal context |
This bundle travels with the claim through every downstream agent. The synthesis agent preserves and merges these mappings — it does not flatten them into prose.
Without this structure, attribution dies the moment any agent summarises. The final report says "studies show X" with no path back to which study, by whom, when. That is not a research output; that is a plausible-sounding hallucination risk.
Two credible sources report different numbers for the same statistic. The wrong moves and the right move:
Annotate with both values and full source attribution. Let the consumer decide.
Renewable energy share of global generation in 2024:
- 30% (IEA, World Energy Outlook 2025, published 2025-10)
- 32% (IRENA, Renewable Capacity Statistics 2025, published 2025-03)
Many apparent contradictions are actually different reporting periods. A 28% figure from 2022 and a 32% figure from 2024 are not conflicting — they show change over time.
| Required in structured output | Why |
|---|---|
| Publication date | Reframes "conflict" as "trend" |
| Data collection date | Distinguishes when data was gathered vs when it was published |
Same logic for: market share, headcount, regulatory status, version numbers, anything that changes.
| Content type | Rendering |
|---|---|
| Financial data | Tables (numbers in prose are hard to scan and compare) |
| News / narrative | Prose (forcing bullets loses causality) |
| Technical findings | Structured lists (multiple discrete items benefit from visual separation) |
Flattening everything into one uniform format loses information. The exam may ask which format suits a given content type — match the content, do not impose a house style.
/compact.When a question describes a failure, ask in order:
Many wrong answers come from reaching for the right idea at the wrong layer:
| Layer | What lives here |
|---|---|
| Prompt content | Case-facts block, structured section headers |
| History management | Summarisation strategy, what to drop, what to keep verbatim |
| Tool integration | Response-shape trimming, structured error envelopes |
| Agent boundaries | Structured key facts and citations crossing agent boundaries — not narrative |
| Escalation policy | Three valid triggers, two unreliable triggers, the frustration nuance |
| Error semantics | Failure type taxonomy, access failure vs valid empty, coverage annotations |
| Session memory | Scratchpad files, subagent delegation, summary injection, /compact, manifests |
| Validation discipline | Stratified sampling, per-field confidence, ground-truth calibration |
| Provenance | Claim-source bundles, conflict annotation, temporal metadata, content-aware rendering |
When you feel the pull toward a fix, ask which layer the question is testing. A summarisation problem is a history-management question (case-facts block), not a prompt-content question ("tell the model to remember the refund amount"). An ambiguous-customer problem is an escalation-policy question (ask for identifier), not a tool-integration question. A novel-error-pattern problem is a validation-discipline question (stratified sampling of high-confidence cases), not a confidence-threshold question. Don't cross layers.
Domain 5's smallest weight is misleading — these concepts surface across other scenarios:
The calibrated-confidence pattern (TS 5.5) and the multi-instance review pattern (Domain 4.6) are the same architectural idea expressed at different layers. Recognise the family resemblance.
Then put the notes to work — drill the scenario questions until the timed exam mode feels easy.