C CCA-F
Take the Test
Last-day revision

Domain cheat sheets.

Condensed revision notes for all five CCA-F domains, structured by exam task statement — because if you can tell which task statement a question is testing, the answer is usually obvious. Global rules first, then the traps, the distractors and the mnemonics. Pass cut: a scaled 720/1000.

✓5 domains, every task statement ✓Distractor cribsheets ✓One-line mnemonics ✓Print it, read it exam morning
DOMAIN 1

Agentic Architecture & Orchestration

27% · highest
Highest single domain weight. Pass cut: 720/1000 scaled.
The rules that decide most questions
  1. Deterministic > probabilistic when stakes are real. Financial / security / compliance / safety = hooks or programmatic gates. Style / tone / formatting = prompts.
  2. Trace failures to their origin. Scope problems = upstream (decomposition / partitioning / context-passing). Quality problems = subagent prompts / tools. Don't blame the agent that faithfully executed a bad brief.
  3. Proportionate fixes. Single agent beats multi-agent when single suffices. Prompts beat hooks when stakes are low. Don't over-engineer.

TS 1.1Agentic Loops

Lifecycle: send → inspect stop_reason → if tool_use, execute tools, append tool_result blocks as role: "user", loop → if end_turn, done.

stop_reason is the ONLY authoritative termination signal.

⚠ Three anti-patterns (all wrong answers)
  • Parsing natural language for "I'm done" phrases
  • Iteration cap as the primary stopping mechanism (caps are safety nets, not termination)
  • Checking content[0].type == "text" — tool_use responses can begin with a text block

Tool result shape

{"role": "user", "content": [
  {"type": "tool_result", "tool_use_id": "<id>", "content": "...", "is_error": False}
]}

Parallel tool_uses → one tool_result per id, all in one user turn.

TS 1.2Multi-Agent Orchestration

Topology: hub-and-spoke only. Coordinator at centre; subagents are spokes; ALL communication routes through the coordinator. Subagents NEVER talk to each other.

The isolation principle (most-misunderstood concept)

  • Subagents do NOT inherit coordinator history
  • Subagents do NOT share memory across invocations
  • Every byte they need must be in their prompt

Coordinator's 7 responsibilities: decompose → select → partition scope → pass context → aggregate → iteratively refine → route for observability.

Failure-tracing rules

SymptomRoot cause
Missing topics in outputDecomposition too narrow
Duplicated work / redundant sourcesScope partitioning sloppy
Shallow on every topicSubagent prompts or tool budgets
Missing source attributionContext-passing destroyed metadata

TS 1.3Subagent Invocation & Context Passing

Task tool is the only mechanism to spawn subagents. allowedTools must include "Task" or the coordinator physically cannot delegate.

AgentDefinition fields: description (used for selection — write like a function docstring), prompt, tools (subset of coordinator's), optionally model.

Context-passing rules

  • Pass complete findings, not pointers
  • Pass structured metadata (claim + source_url + snippet), not collapsed prose
  • Write goal-oriented coordinator prompts, NOT procedural step-by-step

Parallel spawning: emit multiple Task tool_uses in a single coordinator response → executes concurrently. Sequential = separate turns = triple latency. Only safe when subagents are independent.

fork_session: independent branches from a shared expensive baseline. Use for "explore N alternatives from the same analysis."

Decision

  • Many isolated subagents, no shared baseline → parallel Task
  • Branches that share an expensive computed state → fork_session

TS 1.4Workflow Enforcement & Handoff

Enforcement spectrum

PROBABILISTIC ──────────────────────────► DETERMINISTIC
prompt rules → few-shot → routing → hooks / prerequisite gates
~95-99%        ~98%      ~98%       100%

The decision rule (memorise verbatim): If a single failure causes financial loss, security breach, regulatory violation, or safety harm → programmatic enforcement (hooks).

If a single failure causes only a bad-but-recoverable UX → prompts.

Canonical exam scenarios → all answer = hooks: refund verification, trade limits, compliance checks, manager approvals, age/permission gates, audit logging.

⚠ The four-distractor pattern for high-stakes questions
  1. Hook / prerequisite gate ← always correct
  2. Stronger system prompt
  3. Few-shot examples
  4. Specialised subagent with stronger prompt

Multi-concern requests: decompose → investigate in parallel with shared context → synthesise unified resolution.

Human handoff payload (self-contained, no transcript expected)

  • Customer ID
  • Conversation summary (2–3 sentences)
  • Root cause analysis ← exam frequently tests this as missing
  • Action taken so far
  • Recommended action (incl. refund amount) ← exam frequently tests this as missing
  • Compliance / urgency flags

TS 1.5Agent SDK Hooks

HookFiresUse cases
PreToolUseBefore tool executesGates, authorisation, prerequisite checks, financial limits, capturing prior state hash for audit
PostToolUseAfter tool, before model sees resultNormalisation (timestamps, currencies, field names), redaction (PII), truncation, enrichment, result logging

Keyword triggers

  • "Prior state" / "pre-condition" / "before X happens" → PreToolUse
  • "Heterogeneous formats" / "normalise" / "redact" / "model is confused by varying outputs" → PostToolUse

Hooks vs subagents

  • Rule that must be enforced → hook
  • Capability that needs different reasoning / prompt / tools → subagent

Hooks don't rank tools. Hooks gate, block, or transform. Use prompts + tool descriptions for "prefer this tool over that one."

TS 1.6Task Decomposition Strategies

PatternWhenTell
Fixed Sequential PipelineStable structure, each step consumes prior step's output"Every email follows the same flow"
Dynamic Decomposition (typed subagents)Varying requests, role-typed subagents (search/doc/synth)"Same types invoked, scope varies per query"
Orchestrator–WorkerUniform workers differentiated only by prompt-supplied scope"N research workers, each handling a different sub-topic"
Evaluator–OptimiserExplicit checkable quality criteria, first-pass often inadequate"Iterate until coverage / citations / format pass"

Discriminator (Pattern 2 vs Pattern 3): typed roles → dynamic decomposition; uniform clones → orchestrator–worker.

Evaluator–optimiser fails when: the evaluator cannot reliably distinguish good from bad. Subjective criteria → loop spins.

⚠ Over-engineering trap

If a single-agent loop works, the exam's preferred answer is to not refactor to multi-agent. Multi-agent adds latency, cost, and failure surface — warranted only when single-agent genuinely cannot do the job.

TS 1.7Error Handling, Retries, Observability

Three error categories

TypeExamplesResponse
Transient429, 503, network blipRetry with exponential backoff + jitter, capped
Permanent400, 401, 403, 404, malformed inputDon't retry — feed back to model or escalate
LogicalTool returned empty / irrelevant resultsFeed back to model with context, model changes strategy
⚠ Anti-pattern

Blanket retry on every error type. Always classify first.

Idempotency: generate a UUID / hash of {operation + customer + amount} on the first attempt, pass it on every retry, server-side dedup. Without this, a transient retry can double-charge.

Paired-fix pattern (exam favourite): when both retry policy is broken AND duplication occurs, you need both error classification and idempotency keys. Either alone is insufficient.

Errors must reach the model when the model can do something useful. Silent harness error-swallowing (returning empty string for failed tool calls) makes the model produce confident wrong answers.

Observability minimum bar

  • Trace ID per run, span IDs per loop iteration and tool call
  • Structured logging of stop_reason, tool calls, hook decisions, token usage
  • Per-subagent cost & latency telemetry (multi-agent costs compound)
  • Replay capability via stored history

Reliability patterns

  • Circuit breaker — N failures in window → stop calling for cooldown
  • Fallback agent / tool — degraded answer > no answer
  • Cost / token budgets — enforced at harness, not by prompt
  • Per-tool timeouts — no tool blocks the loop indefinitely

CribsheetDistractor Recognition

✗ Reject these for high-stakes scenarios
  • "Add a stronger system prompt instruction…"
  • "Include few-shot examples demonstrating…"
  • "Route to a specialised subagent whose prompt emphasises…"
  • "Increase the model's temperature / use a larger model…"
  • "Forward the full conversation transcript to the human reviewer"
  • "Catch errors and return an empty string silently"
  • "Retry on every error type with the same delay"
  • "Have subagent A pass results directly to subagent B"
✓ Lean toward these
  • "PreToolUse hook that blocks…"
  • "PostToolUse hook that normalises / redacts…"
  • "Idempotency key generated on first attempt"
  • "Pass structured data with claim → source mappings"
  • "Classify errors and handle each category appropriately"
  • "Coordinator's decomposition step is too narrow" (for missing-topic scenarios)
  • "Add Task to coordinator's allowedTools" (for "won't delegate" bugs)
  • "Bounded timeout with fallback message" (for latency hangs)
🧠 One-line mnemonics
  • Loop termination: stop_reason or it's wrong.
  • Subagent state: nothing carries over. The prompt is the universe.
  • High stakes: hooks, every time, no matter how strong the prompt wording.
  • Missing topics: blame the coordinator's decomposition.
  • Duplicate work: blame the coordinator's partitioning.
  • Missing citations: blame the context-passing format.
  • Won't delegate: check allowedTools for "Task".
  • Prior state in audit: PreToolUse only.
  • Heterogeneous tool outputs: PostToolUse normaliser.
  • Double-charged customer: idempotency key.
  • Hanging agent: timeout + fallback, not prompt instruction.
  • Evaluator–optimiser: dies when criteria aren't objectively checkable.
DOMAIN 2

Tool Design & MCP Integration

18%
Appears primarily in: Customer Support Resolution Agent, Multi-Agent Research System, Developer Productivity Tools scenarios.
The rules that decide most questions
  1. Tool descriptions are the routing layer, not documentation. Misrouting is a description problem until proven otherwise. Fix descriptions before adding classifiers, few-shot examples, or merging tools.
  2. Fix at the cheapest layer that addresses the root cause. Edit text before adding infrastructure. tool_choice before prompt instructions. Environment variables before relocating files. Scoped tools before merging agents.
  3. Errors must distinguish "tool couldn't look" from "tool looked and found nothing". isError: true means access failure. isError: false with empty result means success-with-no-data. Conflating them breaks all recovery logic.

TS 2.1Tool Interface Design

The core principle: tool descriptions ARE the routing mechanism. The model has nothing else when choosing between similar tools.

A production-grade tool description contains 5 components

  1. Primary purpose — one sentence
  2. Input contract — formats, types, constraints
  3. Example queries it handles well — concrete strings
  4. Edge cases and limitations — what it won't do
  5. Explicit boundaries vs. similar tools — "Use this for X. For Y, use other_tool." ← the one most teams skip

The misrouting fix-ranking (memorise verbatim)

FixRankWhy
Expand tool descriptions1st — correctLow effort, high leverage, fixes root cause
Few-shot examples in system prompt2ndTreats symptom, token overhead on every call
Merge tools into one3rdLoses useful semantic distinction, high refactor cost
Train a routing classifier4th — wrongExtra model call, latency, new failure point

The exam's preferred answer is ALWAYS "expand descriptions" as the first step for a misrouting scenario.

Tool splitting — the other side

Code smell: a tool description containing "or" between behaviours, or "depending on" between conditions. Example:

"Processes refunds, partial refunds, store credits, or exchanges depending on policy and customer status."

This is unsplittable at runtime. The model can't route to it reliably. Split into purpose-specific tools with tight contracts:

  • issue_full_refund(order_id, reason)
  • issue_partial_refund(order_id, amount, reason)
  • issue_store_credit(customer_id, amount, reason)
  • process_exchange(order_id, replacement_sku, reason)

Renaming a generic tool does NOT fix this. Splitting does.

⚠ The system prompt interaction trap

Even with perfect tool descriptions, misrouting can persist if the system prompt contains keyword-sensitive instructions ("Always look up the customer first") that override tool descriptions. Rule: after fixing tool descriptions, audit the system prompt for keyword conflicts. Don't add more routing rules to the prompt — that competes with the tool descriptions. Remove the conflicting keyword associations instead.

TS 2.2Structured Error Responses

The MCP isError flag is the agent's primary recovery signal. Without structured error metadata, agents guess — and guess badly (retry things that won't succeed, escalate things that needed a moment).

The four error categories (memorise the table)

CategoryExamplesRetryable?Agent response
TransientNetwork timeout, 503, rate limitYes (same input)Backoff + retry, then escalate
ValidationWrong format, missing field, malformed IDYes (after fixing input)Correct input and retry
BusinessRefund exceeds limit, account locked, not eligibleNoSwitch workflow; surface description to user
Permission401, 403, missing scopeNo (by this agent)Escalate or request elevated creds

Structured error payload shape

{
  "isError": true,
  "errorCategory": "business",
  "isRetryable": false,
  "description": "Refund of £450 exceeds £200 self-service limit. Customer must speak with a supervisor."
}

The description for business errors should be customer-friendly — the agent will often surface it to the user.

THE most-tested distinction in 2.2

CaseResponseAgent action
Access failure (tool can't reach source)isError: true, category transient/permissionDecide whether to retry
Valid empty result (tool queried successfully, found nothing)isError: false, empty array / null resultTell user "no results"
⚠ The canonical 2.2 exam trap

Conflating these. Returning [] for both "no such customer" AND "database unreachable" causes the agent to retry endlessly and escalate with a misleading description, when the real answer was "this customer doesn't exist."

Retryability subtlety

  • Transient isRetryable: true = retry with the same input
  • Validation isRetryable: true = retry only after correcting the input

Error propagation in multi-agent systems

Subagents handle local recovery. Only propagate what they cannot resolve.

  • Transient error in a subagent → retry locally, fall back to alternative source, then escalate
  • When propagating, include: what was attempted, partial results if any, recommended action
  • Coordinator should evaluate partial results from successful subagents rather than aborting the whole task on a single subagent failure

Retry budget: even retryable errors have limits. A well-designed tool may include retryAfter and maxRetries hints. "Agent retried 50 times" in a setup = missing retry budget.

TS 2.3Tool Distribution and tool_choice

Tool overload

  • 4–5 tools per agent is the production sweet spot.
  • Reliability degrades sharply past 7–10 tools — softmax over descriptions thins probability margins
  • Scope tools to roles: a synthesis agent should NOT have web_search; a web search agent should NOT have document analysis tools

The tool_choice values (know each verbatim)

Needtool_choice
Model may or may not call a tool"auto" (default)
Model MUST call some tool, of its choice"any"
Model MUST call this specific tool{"type": "tool", "name": "extract_intent"}

Force-a-specific-first-step pattern (exam favourite)

Scenario: "Every customer message must first extract intent and entities before routing. Currently done via system prompt instruction; ~15% of messages skip the extraction."

Correct fix: set tool_choice: {"type": "tool", "name": "extract_intent"} on the first turn. Switch to "auto" for subsequent turns. The model literally cannot output anything other than the forced tool call. 15% miss rate → 0%.

⚠ Distractors to reject
  • "Add stronger instruction to system prompt" — probabilistic, what caused the 15% miss in the first place
  • "Use tool_choice: any" — model could pick a different tool
  • "Validate after the fact" — too late, wrong layer

Rule: when the requirement is "the model MUST do X first", reach for tool_choice: {"type": "tool", "name": "X"}, not a prompt instruction. Prompts request; tool_choice requires.

THE Q9-style pattern — scoped cross-role tools

Setup: synthesis agent frequently returns control to coordinator for one-line fact verifications. 85% are simple lookups, 15% need multi-source cross-referencing. Round-trip latency dominates.

⚠ Wrong fixes
  • Give synthesis agent full web_search + fetch_page → defeats role scoping, tool overload, prompt injection risk
  • Cache → doesn't help with novel claims
  • Merge synthesis + coordinator → destroys architecture
  • Train a smaller model → massive effort for a tool-design problem

Correct fix: add a scoped, constrained verify_fact(claim, max_sources=3) tool to the synthesis agent. Tool description explicitly states that complex multi-source verifications must be routed back to the coordinator.

Pattern in one sentence: for high-frequency simple operations, give the consuming agent a constrained version of the capability; complex cases still escalate to the owner.

Replacing generic tools with constrained alternatives

  • fetch_url(url) → load_document(document_url) — validates allowed domains, document content-types, strips scripts
  • run_sql(query) → query_orders(filters) — single-table, parameterised
  • send_email(...) → send_customer_notification(template_id, ...) — templated, restricted recipients

Same capability, narrower surface, reduced blast radius.

TS 2.4MCP Server Integration

Scoping hierarchy

LevelPathVersion-controlled?Purpose
Project.mcp.json (repo root)Yes (committed, shared)Team-wide integrations for this project
User~/.claude.jsonNo (personal)Personal/experimental servers across all projects

Tool discovery happens at connection time — all tools from all configured servers are made available simultaneously. Too many configured servers re-creates the tool-overload problem at a different layer.

Credential handling — ${VAR} expansion (most-tested 2.4 pattern)

Wrong:

{ "env": { "JIRA_API_TOKEN": "ATATT3xFf_actual_token_here" } }

Right:

{ "env": { "JIRA_API_TOKEN": "${JIRA_API_TOKEN}" } }

Rule: "shared config, personal credentials". The .mcp.json structure stays committed and shared. Secrets live in each developer's local environment.

Don't fix a credential leak by relocating the file. A team integration belongs in .mcp.json so new teammates have it. The fix is parameterisation, not relocation. Also: rotate any leaked token — git history is forever.

MCP resources (content, not actions)

  • Tools = actions the agent can take
  • Resources = content the agent can read (catalogues, schemas, summaries)

Example: jira://projects/INGEST/summary, db://schemas/orders, github://repos/acme/api/issues/labels.

Exam framing: "Agent makes many exploratory tool calls at conversation start to discover available data" → expose resources that catalogue what's available. The agent reads the catalogue once, then queries precisely.

Build vs. use decision

Use a community MCP server when:

  • Standard third-party system (Jira, GitHub, Slack, Notion, Postgres)
  • Maintained server exists
  • Usage pattern fits

Build custom only when:

  • Team-specific internal system
  • Community server fundamentally cannot support the workflow
  • Compliance/security policy prohibits the community server

Exam preference: "Configuration over construction" and "fork over fresh build". When a community server is 90% of what you need (e.g., just want different return fields), post-process its responses or fork it — don't write from scratch.

⚠ The built-in-tool preference trap

When MCP tools overlap with Claude Code built-ins (Grep, Glob, Read), the agent may reach for the built-ins because their descriptions are tighter. Fix: enhance the MCP tool's description with examples and explicit boundary clauses (same principle as 2.1). E.g., "Use this for semantic queries like 'where do we handle payment retries'. Prefer Grep for exact-string searches like 'TODO: fix'."

TS 2.5Built-in Tools

Grep vs Glob — THE most-tested distinction

ToolSearchesReturnsMnemonic
GrepFile contentsMatching lines + filenames"What does the code say?"
GlobFile pathsFile paths matching pattern"What files exist?"

Grep use cases: function callers, error messages, import statements, TODO comments, identifier occurrences.
Glob use cases: all test files (**/*.test.tsx), all configs (**/*.config.{js,ts,json}), files by extension.

⚠ Mirror traps
  • "Find all React component files" → Glob **/*.tsx, NOT Grep import React
  • "Find all callers of formatDate" → Grep formatDate(, NOT Glob (can't match function calls with a path pattern)

Read / Write / Edit

ToolCostWhen
ReadTokens proportional to file sizeWhen you need to see contents
EditLow (diff-only)Targeted modification via unique text anchor
WriteFull file in contextComplete rewrites or Edit fallback

Edit-failure escalation ladder (memorise)

  1. Expand the old_string anchor with surrounding context until unique — the FIRST move
  2. If anchor cannot be made unique (truly identical surroundings) → Read + Write the full file
  3. If intent was actually "change every occurrence" → use replace_all: true

Distractors to reject: "use sed" (wrong tool category), "delete and recreate" (destructive).

Incremental codebase understanding

⚠ Anti-pattern: read everything upfront

Loading ten 500-line files burns tens of thousands of tokens before any reasoning. Context budget destroyed.

Correct pattern:

  1. Grep for entry points (function names, error strings, identifiers)
  2. Read only the entry point file
  3. Grep for the identifiers found in that file across the codebase
  4. Read only the implementations that matter

Wrapper-module tracing: for codebases with re-exports (index.ts does export { formatDate } from './date'), Grep finds both definitions and re-exports. Identify which references are definitions vs. re-exports vs. call sites, then trace the chain.

Grep → Glob sequencing pattern (exam favourite)

Task: "Find all callers of legacyAuth, then find the test files for those callers."

Correct sequence:

  1. Grep legacyAuth( → returns caller files (userService.ts, paymentService.ts, …)
  2. Derive expected test filenames from caller names
  3. Glob **/{userService,paymentService}.test.{ts,tsx} → returns test files

Rule: Grep first when you're discovering content; Glob second to narrow paths around the results. Mirror form: known file set → check contents goes Glob → Grep. The order depends on what is the seed vs. the filter.

CribsheetDistractor Recognition

✗ Reject these
  • "Train a routing classifier model…" (for misrouting — descriptions first)
  • "Merge the two tools into one…" (for misrouting — loses distinction)
  • "Add few-shot examples to the system prompt…" (for misrouting — token cost, treats symptom)
  • "Add a stronger system prompt instruction telling the model to always call X first…" (use tool_choice: {"type": "tool", "name": "X"})
  • "Use a PreToolUse hook to enforce…" (Claude Code harness feature, NOT Anthropic API agent design — wrong layer)
  • "Move the .mcp.json to ~/.claude.json to fix the credential leak…" (use ${VAR} expansion in the shared file instead)
  • "Add .mcp.json to .gitignore…" (defeats team sharing)
  • "Build a custom MCP server because we have specific workflows…" (evaluate community server first)
  • "Give the synthesis agent full web_search…" (defeats role scoping)
  • "Merge the synthesis and coordinator agents…" (destroys architecture)
  • "Increase the retry count…" (for empty-result-misclassified-as-error scenarios)
  • "Glob first to enumerate all source files, then Grep…" (when task is content-seeded)
  • "Read every file in the directory upfront…" (context-budget killer)
  • "Rename the tool to better reflect what it does…" (when the actual problem is splitting)
✓ Lean toward these
  • "Expand both tool descriptions with purpose, input contract, examples, and explicit boundary clauses"
  • "Split the generic tool into purpose-specific tools with tight contracts"
  • "Audit the system prompt for keyword conflicts that override tool descriptions"
  • "Return isError: false with an empty result for the valid-no-data case"
  • "Distinguish access failures from valid empty results in the tool's error contract"
  • "Subagent performs local recovery; coordinator evaluates partial results"
  • "Set tool_choice: {"type": "tool", "name": "X"} to force the mandatory first step"
  • "Add a scoped, constrained verify_fact tool to the consuming agent with explicit fallback to the coordinator for complex cases"
  • "Replace the generic fetch_url with load_document that validates the URL surface"
  • "Use ${VAR} expansion in .mcp.json; rotate the leaked token"
  • "Expose MCP resources for catalogue data so the agent stops probing"
  • "Evaluate the community MCP server first; fork or post-process if a gap exists"
  • "Enhance the MCP tool's description with examples and boundary clauses vs. built-in tools"
  • "Grep for content first, derive filenames, Glob for related paths"
  • "Expand the Edit anchor with surrounding context; fall back to Read + Write only if no unique anchor exists"
🧠 One-line mnemonics
  • Misrouting: descriptions are the routing layer — expand them first.
  • "Or" / "depending on" in a tool description: split the tool.
  • Misrouting persists after fixing descriptions: audit the system prompt for keyword conflicts.
  • isError: true: the tool couldn't look. isError: false + empty: the tool looked and found nothing.
  • Empty array for "no such customer": wrong — that's a valid empty result, not a transient error.
  • Validation retryable: retry only after correcting the input. Transient retryable: retry with the same input.
  • Subagent failure: local recovery first; only propagate what's unresolved, with partial results attached.
  • Coordinator should NOT abort the whole task on one subagent failure if others returned results.
  • Tool count: 4–5 per agent. More degrades selection.
  • "Must do X first": tool_choice: {"type": "tool", "name": "X"} — never a prompt instruction.
  • High-frequency simple operation across roles: scoped cross-role tool with explicit fallback to the owner.
  • Generic powerful tool: replace with constrained alternative (load_document not fetch_url).
  • Credentials in committed .mcp.json: ${VAR} expansion + rotate the leaked token. Don't relocate the file.
  • Exploratory tool calls at conversation start: expose MCP resources.
  • Custom MCP server proposed: evaluate community server first. Fork before fresh build.
  • MCP tool losing to built-in: enhance the MCP tool's description with examples and boundaries.
  • Grep vs Glob: content = Grep, paths = Glob. Wrong choice = wasted time or impossible task.
  • Edit fails with "multiple matches": expand the anchor first; Read + Write only as fallback.
  • Context-budget killer: reading everything upfront. Grep for entry points, Read selectively.
  • Grep → Glob: content-seeded discovery (find usages → find their tests). Glob → Grep: path-seeded check (known file set → check contents).
  • Claude Code hooks (PreToolUse etc.) are CLI-harness features. API-level agent design uses tool_choice, tools, and structured tool results. Don't reach for a hook on an API-design question.

Carry-forwardLayer-Awareness Rule

Many wrong answers come from reaching for the right idea at the wrong layer:

LayerTools available
Anthropic API agent designtool_choice, tools, system prompt, tool result structure, isError
Claude Code CLI harnessPreToolUse, PostToolUse, Stop hooks, settings.json
MCP server config.mcp.json, ~/.claude.json, ${VAR} expansion, resources
Tool implementationError categories, retry hints, input validation, scoped surface

When you feel the pull toward a fix, ask which layer the question is testing. The exam is at the API + MCP + tool-design layers. Hooks are usually the wrong answer in Domain 2.

DOMAIN 3

Claude Code Configuration & Workflows

20%
Appears primarily in: Code Generation with Claude Code, Developer Productivity Tools, Claude Code for CI/CD scenarios.
The rules that decide most questions
  1. Configuration > invocation. When the team needs something to happen consistently, automatically, for everyone — the answer is a file in the right location, not a command someone has to remember to run.
  2. Path prefix is the whole answer. ~/.claude/ is personal, never shared. .claude/ (project) is shared via git. Same filename, opposite scope. Memorise the paths cold; reasoning won't save you.
  3. Load only what's needed. Always-loaded standards live in CLAUDE.md. Pattern-matched standards live in .claude/rules/ with paths: frontmatter. On-demand workflows live in skills/commands. Don't bloat one mechanism with another mechanism's job.

TS 3.1CLAUDE.md Hierarchy

The three levels — memorise the paths verbatim

LevelPathScopeVersion-controlled?
User~/.claude/CLAUDE.mdOnly you, on your machineNo
Project.claude/CLAUDE.md or root CLAUDE.mdEveryone on the repoYes
Directory<subdir>/CLAUDE.mdOnly when working in that subdirYes

Directory CLAUDE.md scopes downward and local, never upward. A src/backend/CLAUDE.md does NOT affect work in src/.

⚠ THE most-tested 3.1 trap — the divergent-developers scenario

Setup: "Developer A's Claude Code follows team conventions. Developer B, on the same repo, same branch, gets inconsistent behaviour."

Root cause 9/10 times: Developer A wrote the conventions to ~/.claude/CLAUDE.md (user-level, lives in their home directory). Git never carried it. Developer B's clone has nothing.

Diagnostic instinct: Divergent behaviour between teammates on the same repo → suspect user-level vs project-level config first.

Trap distractors: "Run /memory reload" — can't refresh a file that doesn't exist on the machine. "Restart Claude Code" — same problem. "Move into directory-level CLAUDE.md" — wrong tool for cross-cutting standards.

Modular organisation (two equally valid patterns)

  1. @import syntax inside CLAUDE.md — reference external files, import per-package standards
  2. .claude/rules/ directory — topic-specific files (testing.md, api-conventions.md, deployment.md) as an alternative to one massive CLAUDE.md

/memory command — the diagnostic tool

Lists which memory files are currently loaded. Use to confirm the diagnosis, not to fix the bug. If a file you expected isn't there, you've found your bug — but the fix is to put the file in the right location, not to run /memory harder.

TS 3.2Custom Slash Commands and Skills

Directory structure — mirrors CLAUDE.md hierarchy

TypeLocationShared?
Project slash commands.claude/commands/Yes — version-controlled
Personal slash commands~/.claude/commands/No
Project skills.claude/skills/<name>/SKILL.mdYes
Personal skills~/.claude/skills/<name>/SKILL.mdNo

Each skill is its own directory containing a SKILL.md file. The directory name becomes the skill name.

Skill frontmatter — the three options to memorise

---
name: brainstorm
description: Explore design alternatives for a feature
context: fork
allowed-tools: [Read, Grep]
argument-hint: <feature-name>
---
OptionEffectExam trigger phrase
context: forkRuns in isolated sub-agent context; verbose output stays in fork; only summary returns to main conversation"produces verbose output", "clutters main conversation", "keep main context clean"
allowed-toolsRestricts which tools the skill can call; prevents destructive actions"analyse only, must not modify", "restrict to read-only"
argument-hintPrompts the developer for required parameters when invoked without arguments"prompt the user for the parameter"

THE critical distinction — Skills vs CLAUDE.md

SkillsCLAUDE.md
When loadedOn-demand, when invokedAlways loaded
PurposeTask-specific workflowsUniversal standards
Example/review, /migrate-component"Use British English in comments"
⚠ Category errors the exam tests
  • "Needs to happen every time" → CLAUDE.md (NOT a skill — skills are opt-in)
  • "Invoked when needed" → skill (NOT CLAUDE.md — would bloat every interaction)
  • "Noisy output, keep main context clean" → skill with context: fork

Personal customisation pattern

Scenario: Team has a /review skill; one developer wants a stricter variant for personal use.

  • Wrong: edit the project skill (affects teammates).
  • Right: create ~/.claude/skills/review-strict/SKILL.md — different name, personal location, no impact on others.

Two independent axes (don't conflate)

  1. Who needs it? → location (.claude/ vs ~/.claude/)
  2. Is the output noisy? → context: fork or not

A skill can be project-scoped AND forked. Personal AND unforked. These are orthogonal.

TS 3.3Path-Specific Rules

Mechanism: files in .claude/rules/ with YAML frontmatter declaring a glob:

---
paths: ["**/*.test.tsx", "**/*.test.ts"]
---

All test files must use describe/it blocks, not test().
Mock external services using shared fixtures in tests/fixtures/.

Loads only when Claude Code is about to work on a file matching the glob. Not loaded for non-matching files.

Why this beats directory-level CLAUDE.md (the comparison the exam tests directly)

Directory CLAUDE.mdPath-specific rules
ScopeOne directory onlyAnywhere the glob matches
Spread across codebaseNeed one CLAUDE.md per directoryOne rule file, globs everywhere
Token costLoaded whenever working in that directoryLoaded only on matching files

Glob patterns match files scattered anywhere. **/*.test.tsx catches every test file regardless of directory. A directory-level CLAUDE.md cannot do this — you'd need N copies that drift.

Token efficiency angle (commonly tested)

  • Path-scoped rules load only when editing matching files → reduces irrelevant context and token usage
  • Project-level CLAUDE.md would burn tokens on every interaction
  • Directory-level CLAUDE.md would require duplicates

When to reach for which mechanism (the mental decision table)

NeedTool
Universal standards (always relevant)Project-level CLAUDE.md
Standards for one specific subdirectoryDirectory-level CLAUDE.md
Standards for files spread across the codebase by patternPath-specific rules
Task-specific workflow, invoked on demandSkill

The signature phrase to recognise: "pattern of files spread across the codebase" → path-specific rules.

Examples that all fit the pattern: **/*.test.*, **/migrations/**/*, **/*.tf, **/api/handlers/**/*.ts.

⚠ Trap distractor

A slash command for the convention. Slash commands require developers to remember to invoke them. Standards must be ambient, not opt-in.

TS 3.4Plan Mode vs Direct Execution

This is a judgement task statement. The exam hands you a task description; you classify.

Plan mode when

  • Large-scale changes (library migration affecting 45+ files, monolith restructure)
  • Multiple valid approaches exist (need to evaluate before committing)
  • Architectural decisions required (service boundaries, abstractions)
  • Multi-file modifications where changes interact
  • Unfamiliar codebase needs exploration before design

Direct execution when

  • Well-understood, narrow scope (single-file bug fix, clear stack trace)
  • Mechanical changes (add a validation conditional, rename a variable)
  • The correct approach is already known

The dividing line: Are there decisions still to make, or just a known thing to do? Decisions → plan mode. Known thing → direct execution.

The Explore subagent

PropertyEffect
Isolates verbose discovery outputSearch results, file listings stay in subagent
Returns summaries to main conversationPreserves main context window
Used during multi-phase tasksPrevents context window exhaustion

Exam phrasing: "main conversation filling up with grep output and file listings" → Explore subagent.

Same principle as context: fork for skills — keep noise out of the main conversation. Different mechanism, same goal.

The combination pattern (commonly tested)

You don't have to pick one mode for the whole task. Common workflow:

  1. Enter plan mode
  2. Explore the codebase, evaluate approaches
  3. Produce a concrete plan
  4. Exit plan mode
  5. Implement the planned steps directly

If a question describes "exploring, then implementing the chosen approach" — that's the hybrid, not a single-mode choice.

Common task classifications

TaskMode
Restructure monolith into microservices (boundaries undecided)Plan mode
Fix null pointer in UserService.getById(), stack trace clearDirect execution
Migrate from winston to pino across 30 files (API differs)Plan mode, then direct execution
Add a date-format validation conditional to one functionDirect execution
React v17 → v18 across 200+ components, concurrent features undecidedPlan mode (then direct execution per file)
⚠ Trap

Calling a multi-file migration "mechanical" when the libraries have different APIs. Scale + interface differences = plan mode.

TS 3.5Iterative Refinement

Technique hierarchy — what to reach for when prose isn't working:

1. Concrete input/output examples (highest leverage)

When prose descriptions get interpreted inconsistently, show 2–3 concrete examples:

Input:  user_name → Output: userName
Input:  api_key   → Output: apiKey
Input:  http_url  → Output: httpUrl

The model generalises from examples more reliably than from descriptions. Three examples eliminate ambiguity that prose admits.

Exam phrasing: "Claude Code interprets the instruction differently each iteration" → concrete input/output examples.

2. Test-driven iteration

Write tests first. Share failures. Claude iterates against a machine-checkable specification rather than your prose. Strongest when behaviour is well-defined but implementation is open.

3. Interview pattern

Have Claude ask questions first before implementing. Surfaces considerations you would miss in unfamiliar domains.

Exam phrasing: "developer working in unfamiliar domain, wants to avoid missing edge cases" → interview pattern.

Batch vs sequence feedback

Issue typeHow to feed back
Fixes interact (changing one affects others)Single message — Claude sees all together, makes consistent decisions
Issues are independent (orthogonal)Sequential — fix one, verify, move on

THE headline rule: when prose fails repeatedly, switch modalities, don't iterate on the prose. Rewording prose tends to produce different misinterpretations, not fewer. Examples constrain the output shape directly.

TS 3.6CI/CD Integration

This is the most memorisation-heavy task statement in the domain. Drill the flags.

The -p flag (print mode) — single most-tested CI detail

  • claude -p "<prompt>" runs in non-interactive mode
  • Without -p, the CI job hangs waiting for interactive input
  • Symptom in exam questions: "the CI pipeline hangs" / "Claude is waiting for input" / "printed banner, then nothing" → answer is -p
⚠ Trap distractors to reject
  • "Increase the timeout" — treats symptom, not cause
  • "Redirect /dev/null to stdin" — wrong layer, not the documented fix
  • "Run in detached shell" — same problem

Structured CI output (for automation)

FlagEffect
--output-format jsonProduces JSON instead of prose
--json-schema <schema>Constrains JSON to a specific shape for reliable downstream parsing

Exam phrasing: "automated system needs to post findings as inline PR comments" → both flags together.

Session context isolation — the counter-intuitive one

The same Claude session that generated code is LESS effective at reviewing its own changes.

It retains the reasoning context that produced the code → less likely to question its own decisions (motivated reasoning).

The exam-correct pattern: use an independent review instance — a fresh Claude Code session with no prior context for the review step.

⚠ Trap

The "efficient" answer feels like "reuse the session, it already knows the code." Wrong. Review needs fresh eyes.

Incremental review context

When a PR gets new commits and you re-run the review:

  • Include prior review findings in context
  • Instruct Claude to report ONLY new or still-unaddressed issues

Without this: developers get duplicate comments on every push → lose trust → start ignoring the bot.

Exam framing: "developers ignoring the bot's comments" / "comment fatigue" / "duplicate findings on every push" → incremental review context, NOT "remove the bot".

CLAUDE.md for CI

CI-invoked Claude Code reads the same .claude/CLAUDE.md. For test generation especially, document:

  • Testing standards (frameworks, patterns)
  • What makes a valuable test in your codebase
  • Available fixtures, mocks, helpers

Without this: CI-generated tests are low-value boilerplate (testing getters, trivial assertions). With it: tests target meaningful behaviour.

Exam phrasing: "CI generates low-quality tests" → strengthen .claude/CLAUDE.md with testing standards.

Test deduplication (commonly missed)

When using CI for test generation, provide existing test files in context alongside the source file. Without them, Claude proposes scenarios already covered — duplicating tests already in the suite.

Two separate context problems in test generation

ProblemFix
Low-quality tests (boilerplate, getters)Document valuable-test criteria and fixtures in .claude/CLAUDE.md
Duplicate tests (scenarios already covered)Provide existing test files in the prompt context

These look similar but have different solutions. The exam may present both options — pick based on which problem is described.

TS 3.6CI/CD Quick Reference — The 5 Failure Patterns

These are the five patterns that most commonly produce wrong answers on CI/CD questions. Memorise the symptom → fix mapping:

SymptomRoot causeFix
CI job hangs indefinitelyClaude Code waiting for interactive inputAdd -p flag
PR comments need to be posted automatically as inline commentsOutput is narrative prose, not structured data--output-format json --json-schema
Duplicate comments on every new commitNo context about what was already reviewedInclude prior findings + instruct "report only new or unaddressed issues"
Review misses issues in code that wasn't changed in the latest commitRunning only the incremental diffStill run full review; prior-findings context handles deduplication
CI generates duplicate testsDoesn't know what's already coveredProvide existing test files in context

What blocks CI most on the exam: the -p flag. If you see "hangs" or "waiting for input" — that is the answer, always.

The independent-session rule appears twice

  • Domain 3 (CI/CD): use a fresh Claude Code session for review, not the one that generated the code
  • Domain 4 (multi-instance review): use an independent Claude instance to review generated code without the generator's reasoning context

Same principle, two domains. On a CI/CD question it maps to -p + separate session. On a prompt-engineering question it maps to multi-instance architecture.

CribsheetDistractor Recognition

✗ Reject these
  • "Run /memory reload to pick up the project's instructions" (when the file doesn't exist on the new dev's machine — it's user-level config that was never shared)
  • "Restart Claude Code to pick up new configuration" (same problem)
  • "Move the file to ~/.claude/CLAUDE.md to make it load on demand" (wrong direction — that removes it from sharing)
  • "Edit the project skill to add stricter behaviour" (when the developer wants personal customisation — create a personal skill instead)
  • "Create a /migrate slash command everyone invokes before writing a migration" (standards must be ambient, not opt-in — use path-specific rules)
  • "Put one CLAUDE.md in every directory that contains test files" (maintenance disaster, drift inevitable — use path-specific rules with ** globs)
  • "Compress the single CLAUDE.md by removing examples and rationale" (compression hides the structural problem — split into .claude/rules/ with paths: frontmatter)
  • "Direct execution — start editing files, adjust as you go" (for unfamiliar 80-file refactor with open architectural questions — plan mode)
  • "Plan mode — even small changes benefit from explicit planning" (for one-line null-check fix with clear stack trace — ceremony, not value)
  • "Rewrite the prose description more carefully" (when prose has failed 3+ times — switch modalities to concrete examples)
  • "Increase the model's temperature" (for prose-misinterpretation problems — examples, not randomness)
  • "Increase the CI timeout to 90 minutes" (for the hanging-CI scenario — -p fixes the root cause)
  • "Redirect /dev/null to stdin" (same)
  • "Switch to --output-format text so output isn't buffered" (the hang isn't an output problem — -p)
  • "Reuse the code-generating session for review — it already knows the context" (motivated reasoning — use a fresh session)
  • "Just remove the review bot, developers find it noisy" (the real fix is incremental review context, not removal)
✓ Lean toward these
  • "The conventions are stored in user-level ~/.claude/CLAUDE.md and were never version-controlled"
  • "Move team standards from ~/.claude/CLAUDE.md to .claude/CLAUDE.md (or root CLAUDE.md) and commit"
  • "Split the monolithic CLAUDE.md into topic files under .claude/rules/ with paths: frontmatter"
  • "Create a personal skill in ~/.claude/skills/ with a different name to avoid affecting teammates"
  • "Add context: fork to the skill so verbose output stays out of the main conversation"
  • "Restrict skill capability with allowed-tools: [Read, Grep] to prevent destructive actions"
  • "Path-specific rule with paths: ["**/*.test.*"] to apply test conventions across the codebase"
  • "Path-specific rule with paths: ["**/migrations/**/*"] for migration conventions everywhere"
  • "Plan mode for investigation and design; direct execution for implementation"
  • "Use the Explore subagent to isolate verbose discovery output from the main conversation"
  • "Provide 2–3 concrete input/output examples of the desired transformation"
  • "Have Claude ask clarifying questions first before implementing (interview pattern)"
  • "Add the -p flag to run non-interactively"
  • "Add --output-format json with --json-schema for machine-parseable findings"
  • "Use a fresh, independent Claude Code session for PR review"
  • "Pass prior review findings into the next review and instruct Claude to report only new or unaddressed issues"
  • "Document testing standards and valuable-test criteria in .claude/CLAUDE.md for CI to read"
🧠 One-line mnemonics
  • Path prefix is the answer. ~/.claude/ = personal. .claude/ = shared.
  • Two devs, same repo, divergent behaviour: user-level config that was never committed.
  • /memory is for diagnosis, not for fix. It tells you what's loaded; it can't conjure a missing file.
  • Directory CLAUDE.md scopes downward and local, never upward.
  • "Or" / "depending on" in a CLAUDE.md section: split into .claude/rules/ with paths: frontmatter.
  • Standards must be ambient. Slash command for a convention = wrong tool. Path-rules or CLAUDE.md.
  • Pattern of files spread across the codebase → path-specific rules with globs.
  • Same conventions in three sibling directory trees → one path-rule with multiple globs, not three CLAUDE.md files.
  • Skill vs CLAUDE.md: on-demand workflow vs always-loaded standard. Don't confuse the categories.
  • Personal customisation of a team skill: new name, ~/.claude/skills/. Never edit the team skill.
  • context: fork for noisy skills. Explore subagent for noisy investigations. Same goal, different mechanism.
  • Plan mode = decisions still to make. Direct execution = known thing to do.
  • Plan mode → direct execution is a valid hybrid, not a wrong answer.
  • Library migration with API differences across N files: plan mode regardless of "mechanical" framing.
  • Prose failed three times: switch to concrete input/output examples. Don't rewrite the prose.
  • Interacting fixes: one message. Independent fixes: sequence them.
  • CI hangs after banner: -p flag. Always.
  • Automated PR comments need machine-parseable output: --output-format json --json-schema.
  • Same session reviewing its own code: motivated reasoning. Use a fresh session.
  • Developers ignoring the review bot: incremental review context (prior findings + "only new issues").
  • Low-quality CI-generated tests: strengthen .claude/CLAUDE.md with test-value criteria and fixtures.

Decision treeConfiguration-Layer Decision Tree

When a question describes a need, ask in order:

  1. Does it need to happen every interaction? → Project-level CLAUDE.md
  2. Does it apply to a pattern of files scattered across the codebase? → Path-specific rule in .claude/rules/ with paths: glob
  3. Does it apply only to one specific subdirectory? → Directory-level CLAUDE.md
  4. Is it a workflow invoked on demand? → Skill in .claude/skills/ (or command in .claude/commands/)
  5. Is it personal-only? → Same answer as 1–4 but in ~/.claude/ instead of .claude/
  6. Is it for the CI to consume? → .claude/CLAUDE.md (still the project-level file — CI reads the same hierarchy)

If the answer to all of 1–4 is "no," you probably don't need configuration — you need a one-off prompt.

Carry-forwardLayer-Awareness Rule

Many wrong answers come from reaching for the right idea at the wrong layer:

LayerWhat lives here
User-level config (~/.claude/)Personal CLAUDE.md, personal commands, personal skills
Project-level config (.claude/)Shared CLAUDE.md, shared commands, shared skills, path-rules, .mcp.json
Directory-level configSubdirectory CLAUDE.md (downward-scoped)
Runtime flags-p, --output-format json, --json-schema
Skill frontmattercontext: fork, allowed-tools, argument-hint
Workflow modePlan mode, direct execution, Explore subagent

When you feel the pull toward a fix, ask which layer the question is testing. A CI hang is a runtime-flag question (-p), not a CLAUDE.md question. A divergent-behaviour bug is a path-prefix question (user vs project), not a /memory question. Don't cross layers.

DOMAIN 4

Prompt Engineering & Structured Output

20%
Appears primarily in: Claude Code for CI/CD, Structured Data Extraction scenarios.
The rules that decide most questions
  1. Specific categorical criteria beat vague confidence-based instructions. "Flag X, report Y, skip Z" with concrete categories beats "be conservative" every time. If an option contains the words conservative, careful, or high-confidence without defining them in categorical terms, it is almost always a distractor.
  2. Match the mechanism to the failure mode. Few-shot for inconsistent judgement. Schema design for fabrication. Validation-retry for output errors. Multi-instance for self-review limitations. Wrong-tool answers are the most common distractor shape in this domain.
  3. Retry fixes output errors. Schema design fixes source gaps. Information genuinely absent from the source document is not recoverable through retry — only through nullable fields, "unclear" enums, or "other" categories. This boundary is tested in nearly every form.

TS 4.1Explicit Criteria

The core principle: Specific categorical criteria obliterate vague confidence-based instructions.

Wrong (vague)Right (categorical)
"Be conservative.""Flag comments only when claimed behaviour contradicts actual code behaviour."
"Only report high-confidence findings.""Report bugs and security vulnerabilities. Skip minor style preferences and local patterns."

Structure to memorise: What to flag. What to report. What to skip. Three categories, no adjectives about confidence.

The false-positive trust problem — heavily tested

When one category has a high false-positive rate, developers stop trusting all categories, including the accurate ones.

The counterintuitive fix: temporarily disable the noisy category entirely while improving its prompt. Don't tighten in place — pull it from output, restore trust in what remains, then re-enable.

⚠ Trap distractors
  • "Lower the confidence threshold for that category" — keeps noise flowing
  • "Add more instructions to the prompt for that category" — keeps noise flowing
  • "Lower the temperature" — deterministic bad criteria is still bad criteria

Severity calibration — the prose-vs-code distinction

Prose definitions of severity drift across runs. Anchor severity levels with actual code examples for each level.

Critical: SQL injection from unvalidated input
  Example: db.query("SELECT * FROM users WHERE id=" + req.params.id)

Medium: Missing error handling on non-critical path
  Example: fs.readFile(path, cb) where cb ignores err parameter
⚠ Trap

A "comprehensive rubric" of five severity levels described in prose. Looks thorough — fails because it lacks code anchors.

TS 4.2Few-Shot Prompting

The headline rule: Few-shot examples are the most effective technique for consistency. Not more instructions. Not confidence thresholds. Not temperature changes. Examples.

Three deployment triggers — memorise verbatim

  1. Detailed instructions alone produce inconsistent formatting
  2. Model makes inconsistent judgement calls on ambiguous cases
  3. Extraction tasks produce empty/null fields for information that exists in the document

Construction rules

PropertyRequirement
Count2–4 targeted examples for ambiguous scenarios (not 10, not 1)
ContentEach example shows the reasoning for why one action was chosen over plausible alternatives
PurposeGeneralisation to novel patterns, not pattern-matching pre-specified cases

The hallucination reduction effect — exam favourite

Documents vary in structure: inline citations vs bibliographies, narrative prose vs structured tables. Few-shot examples covering varied structures dramatically reduce hallucination because the model learns the shape of valid extractions across formats, rather than inventing data to fit the schema.

⚠ Trap distractors
  • "Add more detailed step-by-step instructions covering all edge cases" — instructions grow linearly, edge cases grow combinatorially
  • "Add 10 diverse examples" — wrong number; 2–4 targeted on ambiguous cases beats 10 on solved cases
  • "Lower temperature to 0 and take the union of two runs" — determinism isn't a judgement problem

Signature phrase to recognise: "inconsistent output across runs despite detailed instructions" → few-shot examples.

TS 4.3Structured Output with tool_use

The reliability hierarchy

ApproachGuarantees
Prompt-based JSON ("return JSON like...")Nothing. Model can produce malformed JSON.
tool_use with JSON schemaEliminates syntax errors entirely.

That is the only thing tool_use guarantees. Syntactic validity. Nothing else.

What tool_use does NOT prevent — the three persistent failure modes

  1. Semantic errors — line items don't sum to stated total. Schema doesn't check arithmetic.
  2. Field placement errors — right value, wrong field. Schema validates types, not meaning.
  3. Fabrication — model invents values for required fields when source lacks the information.
⚠ The fabrication factory

A strict schema with all-required fields is a fabrication factory. Required doesn't guarantee correctness — it guarantees something will be present, even if invented.

tool_choice — three modes the exam tests

ValueBehaviourWhen to use
"auto" (default)Model may return text instead of calling a toolWhen the model legitimately might have nothing to extract
"any"Must call a tool, model picks whichGuaranteed structured output, unknown document type, multiple tools available
{"type": "tool", "name": "..."}Must call this specific toolForce a mandatory first step or single-tool workflow

Mental shortcut: named-tool forces a specific tool; "any" forces some tool from the menu.

⚠ Trap

Named-tool offered as a distractor when document type varies. Named-tool routes heterogeneous documents through one tool that doesn't fit most of them.

Schema design — the anti-fabrication toolkit

TechniqueFailure mode it addresses
Optional / nullable fieldsSource legitimately lacks the information → prevents fabrication
"unclear" enum valueAmbiguous cases → lets the model express uncertainty within the schema
"other" + freeform detail stringCategories outside the enum → captures reality instead of forcing wrong-bucket
Format normalisation rules in the promptSchemas validate types; prompts specify formats (e.g., ISO-8601 dates)
⚠ Trap distractors
  • "Use tool_use with all fields required to guarantee complete extractions" — fabrication factory
  • "Use tool_use to ensure line items sum to the total" — schemas don't do arithmetic
  • "Add 'do not fabricate' to the prompt while leaving the field required" — prompt cannot override structural pressure

TS 4.4Validation-Retry Loops

The mechanism — three things sent back on validation failure

  1. The original document
  2. The failed extraction output
  3. The specific validation error message

The model uses the error to self-correct. Dramatically more effective than naive retry, because the model now has signal about what specifically went wrong.

The retry effectiveness boundary — most-tested concept in 4.4

EFFECTIVE forINEFFECTIVE for
Format mismatches (date format wrong)
Structural output errors (object vs array)
Misplaced values (value in wrong field)
Information genuinely absent from the source

Memorise as a sentence: Retry handles output errors. Schema design handles source gaps.

⚠ Trap

A scenario where 15% of documents lack a field and the option says "add validation-retry to recover the missing data." Wrong — the data isn't missing from the extraction, it's missing from the source. Retry will produce fabrications, not corrections. The correct fix is nullable field.

detected_pattern fields — systematic improvement

Add a detected_pattern field to each structured finding (e.g., "unvalidated_user_input_in_sql_concatenation").

Capability unlockedWhy it matters
Analyse dismissal patternsIf pattern X is dismissed 80% of the time, you have data to improve its prompt
Convert black-box tool to measurable systemImprovement based on evidence, not anecdote

Exam framing: "how do you improve a code review bot's accuracy over time" → systematic data path via detected_pattern, not "iterate on the prompt".

Self-correction flows — two patterns

  1. Calculated vs stated totals. Extract stated_total AND calculated_total (sum the line items yourself). Disagreement flags OCR errors, fraud, or model arithmetic mistakes — three things schemas can't catch.
  2. conflict_detected boolean. When source data is internally inconsistent (header date ≠ footer date), surface the conflict rather than silently picking one.

Shared principle: extract redundantly and let inconsistencies bubble up.

⚠ Trap distractors
  • "Add retry-with-feedback to handle all extraction failures" — conflates output errors with source gaps
  • "Increase the retry limit from 3 to 10" — if 3 retries didn't fix it, more produce more fabrications, not more accuracy

TS 4.5Batch Processing

CI/CD note: Scenario 5 (Claude Code for CI) is a primary domain for Domain 4. The batch API and multi-pass review questions in this domain appear in CI/CD framing. The rules below apply directly.

Message Batches API — five facts to memorise

  1. 50% cost savings vs synchronous
  2. Up to 24-hour processing window — no faster, no slower-bounded
  3. No guaranteed latency SLA — could return in 10 minutes or 23 hours
  4. Does NOT support multi-turn tool calling within a single request — single-shot completions only
  5. Uses custom_id for correlating request/response pairs

The fourth point is the sleeper trap. Agentic workflows (model calls tool, sees result, calls another tool) cannot run on batch.

The matching rule — central exam test

Workflow typeAPI choice
Blocking (someone or something is waiting)Synchronous
Latency-tolerant (no human waiting)Batch

Blocking examples: pre-merge CI checks, IDE inline suggestions, customer-facing chat.
Latency-tolerant examples: overnight test generation, weekly compliance audits, monthly competitive intelligence.

⚠ THE Q11-style trap — manager proposes "switch everything to batch to save 50%"

The correct answer is never "migrate everything to batch." It is always the mixed strategy: batch the latency-tolerant workloads, keep blocking workloads synchronous. 50% saving on API cost is dwarfed by the cost of developers blocked waiting for a 6-hour batch return.

Special case — multi-turn tool calls: even if the workflow looks latency-tolerant, multi-turn agentic loops cannot run on batch. Must be synchronous.

Batch failure handling

PatternWhat it does
Identify failures by custom_id, resubmit only failures with modificationsDon't re-batch the whole 100,000; chunk oversized docs and resubmit just the failures
Refine prompts on a sample set BEFORE batch processingIteration loops in batch are ≤24h each; validate synchronously first
⚠ Trap distractors
  • "Use batch for the CI/CD code review bot to save costs" — CI is blocking
  • "Iterate on the prompt during batch runs" — iteration loops are 24h+, validate upfront
  • "Resubmit the entire batch" when only 200 of 50,000 failed — wastes savings

TS 4.6Multi-Instance Review

The self-review limitation — foundational principle: A model reviewing its own output in the same session retains the reasoning context that produced the output. It is biased toward its prior conclusions and less likely to question its own decisions.

The fix: spawn an independent instance — fresh session, no prior context — for the review step. The cold instance catches subtle issues the original misses.

⚠ Trap distractor

"Ask the same model to review its work before returning." Self-review within session. Wrong. Use a fresh instance.

This is the same principle as the CI/CD session-isolation rule in Domain 3.6 — different scenario, same architecture.

Multi-pass architecture for codebase-wide tasks

PassPurpose
Per-file local analysis (one pass per file, focused context)Consistent depth per file
Separate cross-file integration passCatches data flow issues, contract mismatches across files

Why single-pass fails on large repos

  • Attention dilution — early files get more attention than late files (or vice versa) unpredictably
  • Contradictory findings — model loses track of conclusions on file 3 by the time it reaches file 47
  • Misses cross-file issues — cross-file reasoning competes with per-file detail
⚠ Trap distractor

"Use a larger context window so the whole codebase fits in one pass." Context size doesn't fix attention dilution. Bigger window ≠ more attention per file.

Confidence-based routing

StepWhat to do
Model self-reports confidence per findingNumeric or categorical
Route low-confidence findings to human reviewAuto-apply high-confidence findings
Calibrate thresholds using a labelled validation setDon't pick by intuition
⚠ Trap

"Set confidence threshold to 0.9 to ensure quality" without mention of calibration. 0.9 in one system is permissive; in another it's prohibitive. Always calibrate empirically.

CribsheetDistractor Recognition

✗ Reject these
  • "Be more conservative" / "only report high-confidence findings" (vague, unmeasurable — need categorical criteria)
  • "Lower temperature to fix inconsistent judgement" (judgement isn't a randomness problem)
  • "Lower the confidence threshold on the noisy category" (keeps noise flowing — disable and iterate)
  • "Add more detailed instructions to fix inconsistent output" (instructions have ceilings — switch to few-shot)
  • "Add 10 diverse few-shot examples to cover all cases" (wrong number — 2–4 targeted on ambiguous cases)
  • "Use tool_use with all fields required to guarantee complete extractions" (fabrication factory)
  • "Add 'do not fabricate' to the prompt while keeping fields required" (prompt can't override structural pressure)
  • "Use tool_use to ensure line items sum to the stated total" (schemas don't do arithmetic)
  • "Use tool_choice: 'auto' when structured output is required" (lets the model escape into prose)
  • "Force {type: 'tool', name: 'extract_invoice'} for documents of unknown type" (named-tool ≠ heterogeneous routing)
  • "Add validation-retry to recover fields missing from the source" (retry can't conjure absent data)
  • "Increase the retry limit from 3 to 10 to recover difficult cases" (more retries ≠ more accuracy)
  • "Switch everything to the Batch API to save 50%" (breaks blocking workflows)
  • "Batch the CI code review bot" (CI is blocking)
  • "Use batch for the multi-turn customer support agent" (batch doesn't support multi-turn tool calling)
  • "Iterate on the prompt while batch is running" (24h iteration loops)
  • "Resubmit the entire batch when 0.4% fail" (wastes the savings — resubmit failures by custom_id)
  • "Ask the model to review its own findings before returning" (self-review in session — motivated reasoning)
  • "Use a larger context window so the whole codebase fits in one pass" (attention dilution isn't a token problem)
  • "Set confidence threshold to 0.9 to guarantee quality" (without calibration, the number is meaningless)
  • "Increase temperature to encourage broader exploration of issues" (consistency problem, not exploration problem)
✓ Lean toward these
  • "Flag X when condition Y; report Z; skip W" (categorical criteria, not adjectives)
  • "Disable the noisy category temporarily while improving its prompt; keep other categories active" (restores trust)
  • "Add severity-level definitions with concrete code examples for each level" (prose drifts, code anchors)
  • "Add 2–4 few-shot examples covering ambiguous cases, each showing reasoning" (judgement consistency)
  • "Add few-shot examples covering varied document structures (inline citations, bibliographies, narrative, tables)" (hallucination reduction)
  • "Make the field nullable in the schema and instruct the model to return null when not present" (anti-fabrication)
  • "Add an 'unclear' enum value for ambiguous cases" (express uncertainty within the schema)
  • "Add an 'other' enum value plus a freeform detail string" (extensible categorisation)
  • "Switch tool_choice to 'any' so the model must call one of the available tools" (guaranteed structured output across document types)
  • "Force a specific tool with {type: 'tool', name: '...'} to mandate a first step" (single-tool or mandatory-step workflows)
  • "Add validation-retry with error feedback: original document + failed extraction + validation error" (output-error recovery)
  • "Add a detected_pattern field to enable systematic analysis of dismissal patterns" (improvement over time)
  • "Extract calculated_total alongside stated_total and flag discrepancies" (self-correction)
  • "Add a conflict_detected boolean for internally inconsistent source data" (surface ambiguity)
  • "Keep blocking workflows synchronous; batch only the latency-tolerant overnight jobs" (mixed strategy)
  • "Identify failures by custom_id and resubmit only those with modifications" (efficient batch recovery)
  • "Refine prompts on a sample synchronously before batch processing" (avoid 24h iteration loops)
  • "Run extraction in a fresh, independent Claude instance for the review pass" (avoids motivated reasoning)
  • "Run a per-file analysis pass, then a separate cross-file integration pass" (consistent depth + cross-file detection)
  • "Route low-confidence findings to human review, with thresholds calibrated on a labelled validation set" (calibrated, not arbitrary)
🧠 One-line mnemonics
  • Categorical criteria beat confidence adjectives. "Flag X, report Y, skip Z" wins.
  • High false-positives in one category destroy trust in all categories. Disable, iterate, re-enable.
  • Prose severity drifts. Code-example severity anchors.
  • Inconsistent output despite detailed prompts → few-shot, not more instructions.
  • 2–4 targeted examples on ambiguous cases. Not 10. Not 1.
  • Few-shot examples reduce hallucination by showing varied document structures.
  • tool_use eliminates syntax errors. That's it. Not semantics, not placement, not fabrication.
  • All-required schemas are fabrication factories.
  • "auto" for optional. "any" for must-call-something. Named for must-call-this.
  • Nullable fields, "unclear" enums, "other" categories — the anti-fabrication toolkit.
  • Retry fixes output errors. Schema design fixes source gaps.
  • Retry-with-feedback = original + failed output + specific error message.
  • detected_pattern turns a black-box bot into a measurable system.
  • Calculated vs stated total = self-correction without trusting either alone.
  • Batch: 50% cost, ≤24h, no SLA, no multi-turn tool calling, custom_id correlates.
  • Blocking → synchronous. Latency-tolerant → batch. Never "everything to batch."
  • Multi-turn tool calls → synchronous, regardless of latency tolerance.
  • Refine prompts synchronously before batching. Don't iterate inside a 24h loop.
  • Resubmit batch failures by custom_id, not the whole batch.
  • Same session reviewing its own work = motivated reasoning. Use a fresh instance.
  • Per-file pass + cross-file integration pass beats one giant pass.
  • Larger context window doesn't fix attention dilution.
  • Calibrate confidence thresholds on a labelled validation set. Intuition picks the wrong number.

Decision treeMechanism-to-Failure Decision Tree

When a question describes a failure, ask in order:

  1. Inconsistent judgement across runs / inconsistent formatting? → Few-shot examples (2–4, with reasoning)
  2. Malformed JSON? → tool_use with JSON schema
  3. Required field fabricated when source lacks data? → Nullable field in schema + prompt instruction to return null
  4. Ambiguous categorisation? → "unclear" enum value or "other" + freeform detail
  5. Format mismatch in output (e.g., date format wrong)? → Validation-retry with error feedback
  6. Line items don't sum to total / semantic inconsistency? → Self-correction (extract both, compare, flag)
  7. Internally inconsistent source data? → conflict_detected boolean, surface ambiguity
  8. Blocking workflow (someone waiting)? → Synchronous API
  9. Latency-tolerant single-shot workflow? → Batch API
  10. Multi-turn agentic workflow? → Synchronous, always (batch doesn't support it)
  11. Inconsistent depth across files in a large repo? → Per-file pass + cross-file integration pass
  12. Need a second opinion on Claude's work? → Independent instance, fresh session
  13. Want to improve a review bot over time? → detected_pattern field + analysis of dismissal patterns
  14. Bot output dismissed by developers because of noisy category? → Disable that category; iterate; re-enable

Carry-forwardLayer-Awareness Rule

Many wrong answers come from reaching for the right idea at the wrong layer:

LayerWhat lives here
Prompt contentCriteria, severity definitions with code examples, format normalisation rules
Few-shot examplesJudgement consistency, varied-structure generalisation, hallucination reduction
Schema designAnti-fabrication (nullable, unclear, other), required vs optional
tool_choiceWhether and which tool gets called
Validation-retryRecovering output errors
Pipeline architectureMulti-pass review, calculated vs stated, conflict booleans
API selectionSynchronous vs batch based on blocking vs latency-tolerant
Multi-instanceFresh sessions for independent review

When you feel the pull toward a fix, ask which layer the question is testing. A fabrication problem is a schema-design question (nullable), not a prompt-content question ("don't fabricate"). A batch-failure question is a custom_id question, not a retry question. A self-review limitation is a multi-instance question, not a prompt-content question. Don't cross layers.

Cross-domainCI/CD Cross-Domain Map

CI/CD questions span both Domain 3 and Domain 4. When a question is framed as "Claude Code for CI/CD," the answer might live in either domain. Use this map to route quickly:

Problem typeDomainFix
Pipeline hangs / waits for inputD3-p flag
Need machine-parseable PR commentsD3--output-format json --json-schema
Duplicate comments after new commitsD3Prior findings in context + "report only new issues"
CI generates duplicate testsD3Provide existing test files in context
Review generates low-quality testsD3Strengthen .claude/CLAUDE.md with test criteria
Vague instructions don't reduce false positivesD4Categorical criteria: "flag X, skip Y, report Z"
Instructions produce inconsistent outputD42–4 few-shot examples with reasoning
Pre-merge check vs overnight report — which uses batch?D4Pre-merge = sync (blocking). Overnight = batch (latency-tolerant)
Large PR review gives inconsistent depth per fileD4Per-file local passes + separate cross-file integration pass
Generated code reviewed by the generator sessionD4Independent instance (fresh session, no prior context)

The one rule that appears in both domains in different forms: An agent/session that produced output should not review that same output. Use an independent instance. In Domain 3 it's framed as "CI code review session." In Domain 4 it's framed as "self-review limitation." Same answer.

DOMAIN 5

Context Management & Reliability

15%
Smallest weight, but cascades into Domains 1, 2 and 4. Appears primarily in: Customer Support Resolution Agent, Multi-Agent Research System, Structured Data Extraction, Codebase Exploration scenarios.
The rules that decide most questions
  1. Structure beats summarisation. Persistent fact blocks, structured errors, claim-source bundles — anything you need to survive compression must live in structured form, never in summarised prose. The moment a high-value fact becomes free-text, it is at risk of being lost in summarisation, attention dilution, or downstream synthesis.
  2. Explicit signals beat inferred signals. Explicit human request → escalate. Explicit access failure type → distinguish from empty result. Explicit publication date → distinguish trend from contradiction. Answers that rely on the model inferring intent, failure cause, or recency are almost always wrong.
  3. Calibrate, don't trust raw self-reports. Raw model confidence, sentiment, "the model seems sure" — unreliable. Calibrated confidence against a labelled validation set, stratified by document type and field — reliable. This boundary is tested in 5.2 and 5.5.

TS 5.1Context Preservation

The headline rule: Summarisation is lossy compression that disproportionately destroys high-value tokens.

What summarisation compresses away first

Token typeExample loss
Numerical values"$247.83 refund" → "a refund"
Dates"March 3rd" → "recently"
Order/case IDs"#8891" → omitted entirely
Percentages / quantities"32% increase" → "an increase"
Customer-stated expectations"by Friday" → vanished

The fix — persistent "case facts" block

Extract transactional facts into a structured block. Inject verbatim into every prompt. Summarisation never touches it. Conversation history above and below can be compressed; the facts block is sacred.

<case_facts>
- Customer ID: 4471
- Order: #8891, placed 2026-03-03
- Refund requested: $247.83
- Customer deadline: by Friday 2026-05-22
- Channel preference: email
</case_facts>

Lost in the middle — structural attention bias

Transformer attention is empirically biased toward the beginning and end of long inputs. Information buried in the middle is measurably less likely to be retrieved correctly. This is not fixable with a better model — it is structural.

MitigationWhy
Place key findings summaries at the beginningBeginning gets reliable attention
Use explicit section headers (## Findings, ## Sources)Structural anchors aid navigation
Surface relevant slice of long tool results at the top, full payload belowCritical content escapes the middle

Tool result trimming — pre-context filtering

A 40-field order lookup when 5 fields are needed = 35 fields of noise. Trim at the integration layer, before the result hits the context window. Untrimmed results accumulate across turns — turn 8 still carries the noise from turns 1–7.

Define a per-tool response shape with only the fields downstream reasoning needs. Treat tool responses like API contracts, not data dumps.

Full history requirements — stateless API

Claude API calls are stateless. Every request must include the complete conversation history. Drop earlier messages → break coherence. Use prompt caching on the stable prefix (system prompt + tools + early conversation) to make this economical.

Upstream agent optimisation — structured output across agent boundaries

In multi-agent systems, the cheapest place to fix downstream context bloat is upstream. Modify subagents to return:

  • Key facts
  • Citations
  • Relevance scores
  • Confidence values

Not verbose prose, not reasoning chains, not narrative. Reasoning happens inside the subagent's context; only distilled conclusions cross the boundary.

⚠ Trap distractors
  • "Apply progressive summarisation uniformly across the conversation" — destroys transactional spine
  • "Drop the oldest turns to stay under the limit" — breaks coherence; API is stateless
  • "Trust the summarisation prompt to preserve important facts" — it can't; that's the trap
  • "Pass full tool responses to maintain fidelity" — accumulates noise across turns
  • "Return verbose reasoning from subagents so the orchestrator has full context" — fabrication of context bloat

TS 5.2Escalation and Ambiguity Resolution

The headline rule: Three triggers warrant escalation. Two superficially plausible triggers do not.

Three valid escalation triggers

TriggerRequired action
Customer explicitly requests a humanEscalate immediately. Do not attempt resolution first.
Policy exception or gap (request outside documented policy)Escalate — the agent cannot invent policy
Inability to make meaningful progressEscalate after legitimate avenues tried

Two unreliable triggers — wrong answers

Bad triggerWhy it fails
Sentiment-based (customer is frustrated)Frustration does not correlate with case complexity
Self-reported model confidenceModel is often incorrectly confident on hard cases and uncertain on easy ones

The frustration nuance — heavily tested

SituationCorrect action
Straightforward issue + frustrated customer (no human request)Acknowledge frustration, offer resolution, do not escalate yet
Customer reiterates preference for a human after you offer helpNow escalate
Customer explicitly says "I want a human" up frontEscalate immediately — even if the issue looks trivially resolvable

Memorise the distinction: Explicit human request = immediate escalation. Frustration alone = offer help first.

Ambiguous customer matching

Search returns multiple matches for "John Smith." The correct move is always:

Ask for an additional identifier — email, phone, order number, postcode.

⚠ Wrong heuristics — exam distractors
  • Pick the most recent account
  • Pick the most active account
  • Pick the one whose order details "match" the description
  • Pick based on geographical proximity

The cost of asking is one extra turn. The cost of acting on the wrong account is enormous (wrong refund, wrong PII exposure, wrong cancellation). Never resolve ambiguity with heuristics when the customer can disambiguate it in one turn.

⚠ Trap distractors
  • "Resolve the duplicate charge first to demonstrate value, then escalate" — overrides explicit human request
  • "Ask the customer to confirm they really want a human" — patronising; explicit means explicit
  • "Escalate because the customer sounds frustrated" — sentiment is unreliable
  • "Auto-escalate when model confidence drops below 0.6" — uncalibrated self-reported confidence
  • "Pick the most recently active account that matches the search" — heuristic on ambiguity

TS 5.3Error Propagation

The headline rule: How a failure is reported matters as much as that it happened. Silent suppression and workflow termination are both anti-patterns.

Structured error context — four required fields

FieldPurpose
Failure typetransient / validation / business / permission
What was attemptedSpecific query, parameters, endpoint, request shape
Partial resultsAnything gathered before the failure
Potential alternativesSuggested retry, fallback path, narrower query, cached endpoint

Together these give the model enough context to recover. Unstructured errors force the model to give up or hallucinate.

The two anti-patterns

⚠ Both are wrong answers
  • Silent suppression — return {"results": [], "status": "ok"} on tool failure. Orchestrator believes no data exists. Zero recovery path — this is the worst answer in any 5.3 question.
  • Workflow termination — kill the entire pipeline on a single tool failure. Throws away partial results from other subagents that succeeded.

Access failure vs valid empty result — exam favourite

Two superficially identical responses, opposite correct behaviours.

TypeMeaningCorrect response
Access failureTool could not reach the data source (network, auth, timeout)Consider retry or alternative path
Valid empty resultTool reached the source, ran the query, found no matchesNo retry. The empty result is the answer.

The structured error type distinguishes these. Treat all empty results the same and you either retry valid-empties forever or accept access failures as truth. Both wrong.

Coverage annotations — synthesis output discipline

"Section on geothermal energy is limited due to unavailable journal access (3 of 7 expected sources unreachable)."

Better than silently omitting the section. The consumer of the synthesis can now decide whether to act on partial data or wait for the gap to be filled. Silent omission destroys this signal.

⚠ Trap distractors
  • "Return empty results with status=ok so the orchestrator can continue gracefully" — silent suppression, worst option
  • "Kill the entire pipeline on a single subagent timeout" — workflow termination, loses partial results
  • "Retry the tool call indefinitely until it succeeds" — wastes budget on permanent failures; doesn't distinguish failure types
  • "Always retry empty results in case data appears" — conflates valid empty with access failure
  • "Catch the exception silently and log it for later review" — model has no signal during execution

TS 5.4Codebase Exploration

The headline rule: Extended agentic sessions degrade. Verbose discovery output dilutes early findings; the model reverts to generic "typical patterns" instead of session-specific facts.

Two symptoms of context degradation

  1. Model references "typical patterns" instead of the specific classes/functions/files it discovered earlier — generic knowledge crowds out session-specific findings.
  2. Context fills with verbose discovery output (file dumps, search results) and earlier conclusions get buried.

Both symptoms have the same root: high-signal early findings get diluted by high-volume later exploration.

Four mitigation strategies — match the mitigation to the symptom

StrategyBest for
Scratchpad files — write key findings to a file, reference for subsequent questionsFindings you'll need later; survives compaction and crashes
Subagent delegation — spawn a subagent for a specific investigation; coordinator receives distilled result onlyParallel deep dives; keeps coordinator context focused on coordination
Summary injection — summarise findings from one phase before spawning the next phase's subagentsPhase boundaries in multi-phase workflows
/compact — reduce context usage when verbose output fills the windowRecovery valve when nearing the limit; trade-off: loses fidelity

Crash recovery — manifest pattern

Each agent exports structured state to a known file location (a manifest) at meaningful checkpoints. On resume, the coordinator loads the manifest and injects it into the resumed agent's prompts.

The manifest is the source of truth, not the conversation history. Robust to:

  • Process crashes
  • Context window exhaustion mid-task
  • Operator interruptions
  • Long-running jobs spanning multiple sessions

Without a manifest, a crash means starting over. With a manifest, a crash means resuming from the last checkpoint.

⚠ Trap distractors
  • "Increase model temperature to encourage more specific responses" — temperature affects sampling, not retention of specific facts
  • "Re-run the entire exploration from scratch with a fresh context" — wastes work; hits the same problem at the same turn count
  • "Use a larger context window so degradation doesn't occur" — attention dilution isn't a token-count problem (mirrors Domain 4.6 attention rule)
  • "Trust the model to remember earlier findings if asked directly" — that is the failure mode

TS 5.5Human Review and Confidence Calibration

The headline rule: Aggregate metrics hide stratum-level failures. Calibration requires ground truth, not intuition.

The aggregate metrics trap — exam favourite

"Our extraction system achieves 97% accuracy. We can automate."

The trap: 97% overall accuracy can hide a 40% error rate on a specific document type that is rare in the validation set. If that document type is the high-value commercial contracts, your automation is dangerous despite the headline number.

Discipline: Always validate accuracy by document type and by field segment before automating. A single aggregate number is never sufficient evidence to remove humans from the loop.

Stratified random sampling — ongoing audit

Even after calibration and deployment, continue sampling high-confidence extractions for ongoing verification.

ReasonWhy it matters
Novel error patterns emerge over timeNew document templates, edge cases, input distribution drift
Failures slip through silently if not sampledBy definition, they don't trigger existing review thresholds
Stratified across types, fields, confidence bandsEnsures coverage of all segments, not just the loud failures

The discipline: a small, ongoing audit of "the cases we think are fine" is what catches the failures you didn't anticipate.

Field-level confidence calibration — the full pattern

StepDetail
1. Per-field confidence (not per-document)Fields within a document vary wildly in difficulty
2. Calibrate thresholds on a labelled validation setGround truth. Raw model confidence values are not directly meaningful — "0.85" might mean 70% correct or 95% correct depending on field and document type
3. Route low-confidence fields to human reviewThe rest auto-process
4. Prioritise scarce reviewer capacity on highest-uncertainty itemsReviewer time is the limiting resource

Connection to 5.2: This is what makes confidence-based decisions reliable. Naive self-reported confidence is unreliable. Calibrated, field-level, validation-set-grounded confidence is reliable. The exam tests this distinction.

⚠ Trap distractors
  • "97% accuracy meets industry standard, automate confidently" — aggregate metric, no stratum check
  • "Sample only low-confidence extractions for review" — misses novel error patterns in the high-confidence band
  • "Set confidence threshold to 0.9 to ensure quality" — uncalibrated number; meaningless without ground truth
  • "Use document-level confidence to decide review routing" — field variance within documents is the point
  • "Human review is always required regardless of metrics" — too absolute; calibrated automation is valid

TS 5.6Information Provenance

The headline rule: Attribution must travel with every claim as structured data. Conflicts are surfaced, not resolved. Dates distinguish trends from contradictions.

Structured claim-source mappings — five required fields per finding

FieldPurpose
ClaimThe assertion itself
Source URLDirect link
Document nameHuman-readable identifier
Relevant excerptThe actual passage supporting the claim
Publication dateTemporal context

This bundle travels with the claim through every downstream agent. The synthesis agent preserves and merges these mappings — it does not flatten them into prose.

Without this structure, attribution dies the moment any agent summarises. The final report says "studies show X" with no path back to which study, by whom, when. That is not a research output; that is a plausible-sounding hallucination risk.

Conflict handling

Two credible sources report different numbers for the same statistic. The wrong moves and the right move:

⚠ Wrong moves
  • Arbitrarily pick one — erases the disagreement
  • Average the values — fabricates a number that neither source reported — worst kind of synthesis hallucination
  • Pick the more recent without flagging — loses the signal that disagreement existed
  • Pick the source with the "better" reputation silently — erases information the consumer needs
✓ The right move

Annotate with both values and full source attribution. Let the consumer decide.

Renewable energy share of global generation in 2024:
- 30% (IEA, World Energy Outlook 2025, published 2025-10)
- 32% (IRENA, Renewable Capacity Statistics 2025, published 2025-03)

Temporal awareness — date-aware reconciliation

Many apparent contradictions are actually different reporting periods. A 28% figure from 2022 and a 32% figure from 2024 are not conflicting — they show change over time.

Required in structured outputWhy
Publication dateReframes "conflict" as "trend"
Data collection dateDistinguishes when data was gathered vs when it was published

Same logic for: market share, headcount, regulatory status, version numbers, anything that changes.

Content-appropriate rendering — match format to content

Content typeRendering
Financial dataTables (numbers in prose are hard to scan and compare)
News / narrativeProse (forcing bullets loses causality)
Technical findingsStructured lists (multiple discrete items benefit from visual separation)

Flattening everything into one uniform format loses information. The exam may ask which format suits a given content type — match the content, do not impose a house style.

⚠ Trap distractors
  • "Select the more conservative figure when sources conflict" — arbitrary heuristic, erases disagreement
  • "Average the conflicting values to balance the sources" — fabricates a number neither source reported
  • "Omit the statistic when sources disagree" — destroys information
  • "Render everything in a uniform bullet-point format for consistency" — content-format mismatch
  • "Cite sources in a separate appendix rather than inline with claims" — attribution dies during synthesis

CribsheetDistractor Recognition

✗ Reject these
  • "Apply progressive summarisation across the conversation" (destroys transactional spine — use a persistent case-facts block)
  • "Drop the oldest messages to stay under the limit" (breaks coherence; API is stateless)
  • "Pass full tool responses to maintain fidelity" (accumulates noise across turns)
  • "Return verbose subagent reasoning to the orchestrator" (downstream context bloat)
  • "Place the summary at the end of the document" (lost-in-the-middle risk — put it at the beginning)
  • "Resolve the issue first, then escalate if the customer is still unhappy" (when customer explicitly requested a human — overrides explicit request)
  • "Escalate because the customer sounds frustrated" (sentiment ≠ complexity)
  • "Auto-escalate when model confidence drops below threshold X" (uncalibrated self-reported confidence)
  • "Pick the most recently active account when multiple match" (heuristic on ambiguity — ask for an identifier)
  • "Return empty results with status=ok on tool failure" (silent suppression — worst pattern)
  • "Kill the entire pipeline on a single subagent failure" (workflow termination — loses partial results)
  • "Retry empty results in case data appears later" (conflates access failure with valid empty)
  • "Increase temperature to fix context degradation" (temperature isn't the lever)
  • "Use a larger context window to prevent attention dilution" (attention isn't a token-count problem — same as Domain 4.6)
  • "Re-run from scratch with a fresh context window" (hits the same problem at the same turn count — use scratchpad + subagents)
  • "97% aggregate accuracy is sufficient to automate" (masks stratum-level failures)
  • "Sample only low-confidence extractions for review" (misses novel patterns in the high-confidence band)
  • "Set confidence threshold to 0.9 without calibration" (meaningless without ground truth)
  • "Use document-level confidence to route review" (field variance within documents is the point)
  • "Select the more conservative figure when sources conflict" (arbitrary; erases disagreement)
  • "Average conflicting figures to balance the sources" (fabricates a number neither source reported)
  • "Omit the statistic when sources disagree" (destroys information)
  • "Render everything in one uniform format for consistency" (content-format mismatch)
✓ Lean toward these
  • "Extract transactional facts into a persistent case-facts block injected verbatim into every prompt" (anti-summarisation)
  • "Place key findings summaries at the beginning with explicit section headers" (counters lost-in-the-middle)
  • "Trim tool results to relevant fields before appending to context" (prevents noise accumulation)
  • "Include the complete conversation history in every API request" (stateless API)
  • "Modify upstream subagents to return structured key facts and citations, not verbose reasoning" (downstream optimisation)
  • "Escalate immediately when the customer explicitly requests a human, without attempting resolution first" (explicit signal honoured)
  • "Acknowledge frustration, offer resolution; escalate only if the customer reiterates the human request" (frustration nuance)
  • "Ask the customer for an additional identifier (email, phone, order number) to disambiguate" (no heuristics on ambiguity)
  • "Return a structured error: failure type, query attempted, partial results, suggested alternative" (recovery-capable error)
  • "Preserve partial results from succeeded subagents when one fails; annotate the synthesis with a coverage note" (graceful degradation)
  • "Distinguish access failure (consider retry) from valid empty result (the answer is 'none')" (exam favourite)
  • "Write key findings to a scratchpad file and reference it for subsequent questions" (context-degradation fix)
  • "Delegate specific investigations to subagents; coordinator receives distilled result only" (context isolation)
  • "Inject phase summaries before spawning next-phase subagents" (continuity without full history)
  • "Export structured state to a manifest file at checkpoints for crash recovery" (resume from manifest, not history)
  • "Validate accuracy by document type and field segment before automating" (defeats aggregate-metrics trap)
  • "Continue stratified sampling of high-confidence extractions post-deployment to catch novel error patterns" (ongoing audit)
  • "Output per-field confidence, calibrate thresholds on a labelled validation set, route low-confidence to human review" (full calibration pattern)
  • "Prioritise scarce reviewer capacity on highest-uncertainty items" (resource allocation)
  • "Attach structured claim-source mappings (claim + URL + doc + excerpt + date) to every finding" (provenance survives synthesis)
  • "Surface conflicting sources with both values and full attribution; let the consumer decide" (conflict handling)
  • "Include publication and data collection dates in structured outputs to distinguish trends from contradictions" (temporal awareness)
  • "Match rendering to content: tables for financial data, prose for narrative, structured lists for technical findings" (content-format match)
🧠 One-line mnemonics
  • Summarisation compresses numbers, dates, IDs, percentages, expectations first. Extract them into a persistent block that summarisation never touches.
  • Lost in the middle is structural, not a model bug. Put summaries at the beginning, use section headers, never bury the lede.
  • Trim tool results before they hit context. Noise accumulates across turns.
  • The API is stateless. You manage history. Drop messages → break coherence.
  • Upstream subagents return structured key facts. Not narrative. Not reasoning chains.
  • Three valid escalation triggers: explicit human request, policy gap, no meaningful progress.
  • Two unreliable triggers: sentiment, self-reported confidence.
  • Explicit "I want a human" = escalate now, do not investigate.
  • Frustration alone = offer help first; escalate if reiterated.
  • Ambiguous customer match = ask for an identifier, never heuristic.
  • Structured error = type + attempt + partial results + alternatives.
  • Silent suppression is the worst pattern. Empty results with status=ok = no recovery path.
  • Workflow termination throws away partial results. Graceful degradation preserves them.
  • Access failure ≠ valid empty. Retry only the former.
  • Coverage annotations beat silent omission.
  • Long sessions degrade: scratchpad, subagents, summary injection, /compact.
  • Manifest files are the source of truth across crashes. Not conversation history.
  • 97% accuracy can hide 40% on a stratum. Validate by document type AND field segment.
  • Stratified sampling continues post-deployment. Sample high-confidence cases to find novel errors.
  • Per-field confidence, calibrated on a labelled validation set. Raw 0.9 means nothing without ground truth.
  • Every finding carries: claim + URL + doc + excerpt + date. Attribution dies otherwise.
  • Surface conflicts with both values and attribution. Never average. Never silently pick.
  • Dates distinguish trends from contradictions. 28% (2022) vs 32% (2024) is not a conflict.
  • Match rendering to content. Tables for numbers, prose for narrative, lists for technical.

Decision treeMechanism-to-Failure Decision Tree

When a question describes a failure, ask in order:

  1. Conversation history hitting the limit and you fear losing transactional facts? → Persistent case-facts block injected verbatim, summarise surrounding turns only
  2. Key findings buried in long inputs and being missed? → Place summaries at the beginning + explicit section headers
  3. Tool result has 40 fields when you need 5? → Trim before appending to context
  4. Downstream agents drowning in upstream verbosity? → Modify upstream agents to return structured key facts only
  5. Customer explicitly asks for a human? → Escalate immediately, no investigation
  6. Customer is frustrated but issue is simple and no human requested? → Acknowledge, offer resolution, escalate only if reiterated
  7. Multiple customers match the search query? → Ask for an additional identifier (email/phone/order ID)
  8. Tool call failed and you're deciding how to report it? → Structured error: type + attempt + partial results + alternatives
  9. Tool returned an empty array — should you retry? → Distinguish access failure (retry) from valid empty (do not retry; that is the answer)
  10. One subagent fails in a pipeline of four? → Preserve the three successes, annotate the gap, do not terminate the pipeline
  11. Long codebase exploration session: model referencing "typical patterns" instead of specific findings? → Scratchpad file with key findings, referenced for subsequent questions
  12. Need parallel deep dives without bloating the coordinator's context? → Subagent delegation, coordinator receives distilled output only
  13. Long-running job risks a mid-task crash? → Manifest file checkpoint pattern, resume from manifest
  14. 97% accuracy claimed and team wants to automate? → Demand stratified validation by document type and field segment first
  15. Need to catch novel error patterns after deployment? → Stratified sampling of high-confidence extractions, ongoing
  16. Routing extractions to human review? → Per-field confidence, calibrated on labelled validation set, route low-confidence
  17. Two credible sources report different numbers? → Annotate both with full source attribution; let the consumer decide
  18. Same statistic with different values from different years? → Include publication dates; this is a trend, not a contradiction
  19. Rendering the synthesis output? → Match format to content type (table/prose/list)

Carry-forwardLayer-Awareness Rule

Many wrong answers come from reaching for the right idea at the wrong layer:

LayerWhat lives here
Prompt contentCase-facts block, structured section headers
History managementSummarisation strategy, what to drop, what to keep verbatim
Tool integrationResponse-shape trimming, structured error envelopes
Agent boundariesStructured key facts and citations crossing agent boundaries — not narrative
Escalation policyThree valid triggers, two unreliable triggers, the frustration nuance
Error semanticsFailure type taxonomy, access failure vs valid empty, coverage annotations
Session memoryScratchpad files, subagent delegation, summary injection, /compact, manifests
Validation disciplineStratified sampling, per-field confidence, ground-truth calibration
ProvenanceClaim-source bundles, conflict annotation, temporal metadata, content-aware rendering

When you feel the pull toward a fix, ask which layer the question is testing. A summarisation problem is a history-management question (case-facts block), not a prompt-content question ("tell the model to remember the refund amount"). An ambiguous-customer problem is an escalation-policy question (ask for identifier), not a tool-integration question. A novel-error-pattern problem is a validation-discipline question (stratified sampling of high-confidence cases), not a confidence-threshold question. Don't cross layers.

Cross-domainCross-Domain Connections

Domain 5's smallest weight is misleading — these concepts surface across other scenarios:

  • TS 5.1 → Domain 1 (Customer Support Resolution): persistent case-facts block is the canonical fix for multi-turn refund/support cases
  • TS 5.2 → Domain 1: escalation logic owns half of the support-agent question space
  • TS 5.3 → Domain 2 (Multi-Agent Research): structured error propagation is what keeps research pipelines recoverable
  • TS 5.4 → Domain 3 (Codebase Exploration): scratchpad/subagent/manifest patterns are the codebase-agent toolkit
  • TS 5.5 → Domain 4 (Structured Data Extraction): calibration + stratified sampling overlap with Domain 4.6 confidence-based routing
  • TS 5.6 → Domain 2: provenance is the difference between a research synthesis and a plausible-sounding hallucination

The calibrated-confidence pattern (TS 5.5) and the multi-instance review pattern (Domain 4.6) are the same architectural idea expressed at different layers. Recognise the family resemblance.

Print this. Read it the morning of the exam.

Then put the notes to work — drill the scenario questions until the timed exam mode feels easy.

Test Yourself Official Sample Questions