Passing CCAR-F

Brody Networks

An independent, unofficial study companion for the Claude Certified Architect, Foundations exam. Not affiliated with, sponsored by, or endorsed by Anthropic. brodynetworks.substack.com

  1. Sep 4

    Episode 12. Exam Day, and Working the Questions

    The finale. The exam-day experience, pacing sixty items across a hundred and twenty minutes, the five patterns that decide most scenario questions, and three walkthroughs built from the exam guide’s own publicly released sample questions, the only pre-existing items this series ever discusses. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. This finale discusses the exam guide’s own published sample questions from guide section 9, which Anthropic released publicly for exactly this purpose; every other episode’s practice questions were written for this show against the published guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 Exam day, the last episode * 0:54 The confidentiality agreement * 2:16 Rules of conduct and workspace * 2:50 Pacing: sixty items, two minutes each * 3:41 The five recurring patterns * 3:52 Pattern 1: deterministic beats probabilistic * 4:29 Pattern 2: fix the cause at its own layer * 5:04 Pattern 3: the proportionate first step * 5:44 Pattern 4: least privilege * 6:16 Pattern 5: structured data over prose * 6:44 Three walkthroughs from the guide’s own questions * 7:10 Walkthrough 1: identity verification, domain 1 * 8:12 Walkthrough 2: scattered test files, domain 3 * 9:11 Walkthrough 3: batch versus blocking, domain 4 * 10:10 Why the same five patterns keep appearing * 10:37 Retakes, renewal, and criterion-referenced scoring * 11:40 The close Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 9, published sample questions; sections 10, 12, 13, 14, 15) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. Eleven episodes ago, this show was a hunch that a foundations exam could be studied for out loud. Today it is exam day. The last episode. The only one where we talk about actual questions. These particular ones were published by Anthropic itself, in the open, for exactly this purpose. No task statements today. This episode is the room, the clock, and the patterns that decide most of what you will see once you sit down. Start with the room, since it is the part nobody studies for. You will accept a confidentiality agreement before the exam begins, covering every question, answer option, and scenario as Anthropic’s confidential property. Decline it, and the session ends with no refund. Accept it, and from that moment forward, this show and every other honest source goes quiet on your specific paper, forever. You may be tempted, weeks later, to compare notes with someone else who has taken it. Do not. That conversation ends one of two credentials, and it is never worth the trade. That silence is not a policy inconvenience. It is why every question in this series was built from the published guide instead of a real one. And why it had to stay that way. There was never a version of this show that could wait until after someone sat the exam and still call itself independent. The scripting had to happen before, on the public record alone, or not at all. Not a rumor, not a forum post, not something somebody who sat the exam once passed along secondhand. Every fact in every one of these twelve episodes traces back to something Anthropic put on a public page. That is not a limitation this show worked around. It is the entire premise it was built on. Online or at a test center, the rules of conduct are the same in spirit. A clear workspace. No phones, no notes, no second monitor. Webcam in view the whole session if you are testing online. None of this is unusual for a proctored exam. Treat it as background noise, not as something to worry about the morning of. Every certification with real value runs this way. Familiar rules are not a warning sign. They are just the cost of a credential that means the same thing for everyone who holds it. Pacing is arithmetic, and it is worth doing once before you walk in. Sixty items, one hundred twenty minutes. Two minutes a item, on average, and averages are the operative word. Some items you will read once and know. Some will take four minutes of actually working through a scenario. Do not panic at minute forty because you have only finished fifteen items. Fifteen at minute forty is on pace, not behind, once you account for the reading. Panic itself costs more time than the pace ever did. The scenario-based items are front loaded with reading, and that reading is where the two minutes goes. Reading the scenario is not time lost before the real work starts. It is the real work. Most of what decides your answer is sitting in the setup, not in the four options underneath it. Now the five patterns. Not exam content. Just what eleven episodes of task statements add up to, once you stand back from all thirty of them at once. Pattern one: deterministic beats probabilistic, whenever consequences are financial or carry safety weight. A hook or a gate wins over a better sentence. Every single time the downside of a miss is real money or real harm. You heard this in episode three with the refund agent, and it is the most reliable single lens in the whole exam. Ask yourself one question before anything else: if this specific step gets skipped, does anyone lose money or get hurt. If the answer is yes, stop looking for the well written instruction. Look for the gate. Pattern two: fix the cause at its own layer, never compensate downstream. A narrow coordinator decomposition gets fixed at the coordinator, not by adding a fact-checking pass after synthesis. A missing tool description gets fixed by rewriting the description, not by adding a routing layer on top of it. Say the exam offers you a fix that sits one layer away from where the actual problem lives. That fix is almost always the trap. It is the answer that sounds responsible and touches nothing that actually broke. Pattern three: the proportionate first step. Cheap, high leverage fixes before new infrastructure. Expand a tool description before you build a classifier. Add explicit criteria before you deploy a second model to check the first one’s work. The exam consistently rewards the smallest change that actually reaches the root cause. It punishes reaching for machinery before you have tried the obvious thing. A second model, a new pipeline stage, a whole new subagent. All of these are real tools. All of them are wrong the moment a one line fix was sitting right there unused. Pattern four: least privilege, applied to tools and to context both. Scope each agent to what its role actually needs. Scope what a subagent receives to what it actually has to reason about. Every unnecessary tool is a temptation waiting to be misused. Every unnecessary block of context is competing for attention against the parts that matter. Scope is not caution for its own sake. It is removing a failure mode before it ever gets the chance to happen. Pattern five: structured data beats prose at every handoff. Between subagents, into a schema, across a summarization boundary, in an error response. Anywhere information crosses a seam, the version that survives is the one with fields, not the one trusted to a sentence. A well written sentence still has to be re-read and re-interpreted by whatever receives it. A field just has to be looked up. Hold those five, and you can work through unfamiliar scenarios by asking which one applies, rather than hunting for a memorized fact. Let’s prove it, on three of the guide’s own published sample questions. These are not questions written for this show. They are printed in the exam guide itself, publicly released. That is exactly why we can discuss them here, and nowhere else in the series. First, domain one. A support agent skips its identity verification step. It calls the order lookup using only a customer’s stated name, twelve percent of the time. It occasionally misidentifies accounts and issues refunds to the wrong person. The guide asks what change would most effectively fix this. The published answer is a programmatic prerequisite. It blocks the lookup and the refund calls until identity verification has actually returned a confirmed ID. A stronger system prompt and better few shot examples are both offered as tempting alternatives. Both are wrong for the same reason. They are still probabilistic, and this is a financial consequence problem. Pattern one, straight through the middle. Read that scenario again and the tell was there from the first sentence. The failure has a dollar figure attached to it. The moment a scenario names a financial consequence, a gate has already won before you finish reading the four options. Second, domain three. A codebase has react components, a p i handlers, and database models, each with their own conventions. Test files sit scattered next to the code they test, rather than gathered in one place. The guide asks how to make sure Claude Code applies the right convention automatically regardless of location. The published answer is a rules directory with YAML frontmatter glob patterns, matching files by name rather than by folder. A root CLAUDE.md relying on inference is offered, and so is a directory level file that cannot reach scattered files. Both fall short for the reason episode six spent fifteen minutes on. A glob does not care where a file lives. Only what it is. That is the whole difference between a directory-shaped solution and a pattern-shaped one, and it is worth carrying past this exam entirely. Third, domain four. A team wants to cut API costs. Move two workflows onto the Message Batches API for its cost savings. A blocking check that has to finish before a developer can merge. An overnigh

    Episode 12. Exam Day, and Working the Questions
  2. Sep 4

    Episode 11. Reliability, Escalation, Errors, Review, Provenance

    Domain 5, task statements 5.2, 5.3, 5.5 and 5.6. Fifty five percent resolution against an eighty percent target, because an agent escalated the easy cases and attempted the hard ones. The three real escalation triggers, structured error context across multi-agent systems, why an aggregate accuracy number can hide a real problem, and keeping a claim attached to its source through synthesis. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:22 Fifty five percent against an eighty percent target * 1:01 The three legitimate escalation triggers * 1:49 Why sentiment and confidence fail as proxies * 2:31 Honoring an explicit request immediately * 3:07 Escalating on a genuine policy gap * 3:28 Multiple matches: ask, do not guess * 3:58 Structured error context across agents * 4:06 Access failure versus a valid empty result * 4:53 The two anti-patterns * 5:40 Why an aggregate accuracy number can hide a problem * 6:00 Stratified sampling of high-confidence extractions * 6:45 Field-level confidence calibration * 6:50 Provenance and claim-source mappings * 7:38 Annotating conflicting statistics * 7:47 Dates against false contradictions * 8:47 Worked question * 9:35 Why the other three options are there * 10:10 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 5.2, 5.3, 5.5 and 5.6) Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. Fifty five percent resolution against an eighty percent target. Not because the agent was attempting cases it could not handle. It was escalating the easy ones and attempting the hard ones. Almost exactly backwards. Nobody had told it what escalation was actually for. Domain five continues. Task statements five point two, five point three, five point five and five point six. Escalation, error propagation across agents, honest measurement, and keeping a claim attached to its source. The last stretch of the exam blueprint, and one of the highest density episodes in the series. Start with escalation, because most systems get the trigger wrong before they get anything else wrong. There are exactly three legitimate reasons to escalate. The customer explicitly asks for a human. Policy is genuinely silent or ambiguous on the specific request in front of you. The agent cannot make meaningful progress, not that it does not want to, that it actually cannot. Notice what is missing from that list. Complexity is not on it. A case being hard is not, by itself, a reason to hand it to a human. The system I opened with had inverted this exactly. It was escalating on some proxy for difficulty, and treating everything else as safe to attempt. Difficulty was never the right signal in the first place. Sentiment and self reported confidence are the two proxies people reach for anyway. The exam wants you to know precisely why both fail. An upset customer with a completely straightforward request does not need a human. Confidence describes how sure the model feels, and a model can feel very sure while being completely wrong. Neither number measures the thing you actually care about, which is whether this specific case matches one of the three real triggers. A calm customer with an unanswerable request needs a human just as much as a furious one does. Tone tells you almost nothing about which of the three triggers actually applies. There is a nuance inside the first trigger worth holding onto. A customer who explicitly demands a human gets one immediately, no investigation first, full stop. A customer who is upset but has not demanded escalation gets acknowledgment and an offer to resolve. It escalates only if they say so again. Read frustration as an automatic escalation trigger, and you hand off cases the agent could have closed in one turn. Ignore an explicit request for a human, and you have ignored the one trigger that needs no further judgment at all. The policy gap trigger deserves its own moment, because it is not the same as a hard question. Competitor price matching, when your policy only ever addresses your own site’s price adjustments, is not a difficult case. It is a case with no answer written down anywhere. The correct move is escalating precisely because policy is silent. Not attempting a guess and hoping it lines up with whatever the company would have wanted. One more pattern worth naming: multiple matches on a lookup. Two customers with the same name, three orders matching a partial description. The fix is asking for another identifier, not picking the most likely match and hoping. A heuristic guess here is a guess about someone’s account. Getting it wrong is not a near miss. It is the wrong customer’s data in the wrong conversation. One extra question costs thirty seconds. A wrong match costs someone else’s private information landing in a stranger’s conversation. Now task statement five point three. This is domain two’s error work, one level up. Across a multi agent system, instead of inside one tool. Structured error context, failure type, what was actually attempted, partial results, alternatives. That is what lets a coordinator make an actual recovery decision, instead of guessing blind. The same access-failure-versus-empty-result line from a few episodes back matters even more here. A coordinator sitting above several subagents cannot afford to confuse the two, even once. A subagent reporting a timeout needs a different coordinator response than a subagent reporting a genuinely empty, successful search. Collapse those into one undifferentiated failure signal, and the coordinator loses the information it needs to respond correctly to either one. A retry makes sense for one. A retry is wasted effort for the other. From the coordinator’s seat, undifferentiated failure looks identical either way, until something tells it otherwise. Generic error statuses hide exactly this distinction. Search unavailable tells the coordinator nothing about what to do next. Two anti patterns bracket the correct behavior on either side. Silently swallowing an error and reporting empty results as if the query simply found nothing is one. Killing an entire multi agent workflow because one subagent failed is the other. The right answer sits between them. Local recovery inside the subagent for anything transient. Structured context upward for anything that could not be resolved locally. The overall workflow continues on partial results, with the gap honestly annotated, rather than hidden or treated as fatal. Task statement five point five is where honest measurement lives. It opens with a number that sounds reassuring, and is not telling you what you think. Ninety seven percent overall accuracy can hide a document type performing at sixty percent. Buried inside an average that never breaks the number down by segment. Stratified random sampling is the fix. It targets exactly the extractions a naive process would never think to check. The high confidence ones. Confidence is the model’s own self assessment, and self assessment is not ground truth. Sampling across confidence levels, not just auditing the cases the model already flagged as uncertain. That is what actually catches an error pattern the model does not know it is making. Analyzing accuracy by document type and by field, before you reduce human review on the strength of one aggregate number. This is not optional caution. It is the only way to know whether that ninety seven percent is real everywhere. Or real on average and false somewhere specific. An aggregate number is an average of many truths and a few failures. The average alone cannot tell you which document type is hiding the failures. Calibrate field level confidence against labeled validation sets, and route by that calibration. Low confidence and contradictory source documents both belong in front of a human, prioritized by the reviewer time you actually have. Reviewer hours are finite. Spend them where the model is actually unsure, not spread evenly across a pile that mostly did not need a second look. Last task statement: five point six, provenance. When multiple sources feed one synthesis, source attribution is exactly what summarization steps lose first. The same way task statement five point one’s numbers get lost first in a single conversation. A claim survives. Which source said it does not. Not unless something forces it to travel alongside the claim structurally, rather than trusting prose to carry it. The fix is a structured claim-source mapping the synthesis agent is required to preserve, not paraphrase, all the way through. When two credible sources disagree on a statistic, the answer is never picking one and moving on. Annotate the conflict. Attribute both values to their sources. Let whoever reads the report see the disagreement, rather than a single number that quietly erased it. Dates matter here in a specific way. Two sources reporting different numbers six months apart are not necessarily contradicting each other. They may simply be describing different moments. Requiring a publication or collection date on every structured finding is what lets a reader, or a coordinator, tell a real contradiction apart. From an honest change over time. One closing point on how a synthesis actually presents what it found: match the shape of the content to the content itself. Financial data as a table. News as prose. Technical findings as a structured list. F

    Episode 11. Reliability, Escalation, Errors, Review, Provenance
  3. Sep 4

    Episode 10. Context That Survives

    Domain 5, task statements 5.1 and 5.4. An agent quietly loses a customer’s exact refund amount to progressive summarization. What summarization eats first, the lost-in-the-middle effect, trimming tool output at the source, and keeping a long codebase exploration from degrading into generic answers. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:21 The refund amount that got summarized away * 0:59 What summarization eats first * 1:52 The lost-in-the-middle effect * 2:30 Tool output outpacing relevance * 3:14 The case facts block * 3:51 Ordering aggregated inputs * 4:22 Trimming at the source * 4:52 Passing complete conversation history * 5:26 Structured facts over reasoning chains * 6:03 Context degradation in long exploration * 6:56 Scratchpad files * 7:17 Delegating discovery to a subagent * 7:50 Summarizing between exploration phases * 8:20 Structured state manifests for crash recovery * 8:54 /compact * 9:08 Worked question * 10:20 Why the other three options are there * 11:12 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 5.1 and 5.4) * Context windows Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The agent forgot the refund amount from turn three. Not the whole conversation. Just the one number that actually mattered, the three hundred and forty dollars the customer had already been promised. By turn nine, it was offering a different figure, confidently. Somewhere along the way, that first number got summarized down to about three hundred dollars. About is not a number you can issue a refund against. Nobody rewrote that number on purpose. It just got compressed, one summary at a time. Until the version left standing was close enough to sound right, and wrong enough to matter. Smaller than the others, and no easier to lose points on, because reliability failures are the ones customers actually notice. Domain five, context management and reliability, fifteen percent of the exam. Task statements five point one and five point four today. What summarization actually eats first, and how to keep a long exploration from quietly losing the plot. Start with what gets lost, because it is not random. Progressive summarization is good at compressing prose and terrible at protecting numbers. Amounts, percentages, dates, a specific thing the customer said they expected. Those get rounded, softened, and generalized every time a summary compresses them. Three hundred and forty quietly becomes about three hundred. That is not a rounding error. It is a different number. There is a second failure sitting right next to it, and it is about position, not compression. Long inputs get read unevenly. The beginning gets attention. The end gets attention. The middle is where things go missing, not because the model cannot see it, but because it reliably weights it less. A key finding buried in the middle of a long aggregated report is more likely to get skipped. Statistically, compared to the same finding sitting first or last. Nothing about the finding changed. Only where you put it did. That is the entire fix, and it costs nothing more than deciding where a sentence goes. Third failure, and this one is just arithmetic: tool output piles up faster than its relevance does. An order lookup can return forty or more fields. Maybe five of them matter for the question actually in front of you. Every one of those unused fields still sits in context. It still costs tokens. It still competes for attention against the five that actually count. Thirty five fields nobody asked about are not free just because nobody reads them. They are still there, still competing, on every single turn. Multiply that by every order lookup in a long session. The waste compounds far faster than anyone tracking it turn by turn would expect. The fix for the first failure has a name worth knowing: a case facts block. Pull the transactional numbers, amounts, dates, order numbers, statuses, out of the narrative entirely. Hold them in a structured block, included in every prompt, sitting outside whatever gets summarized. Summarization can compress the story. It should never get anywhere near the three hundred and forty dollars. Let the narrative get shorter every turn if it has to. The number does not get to shrink with it. Everything else in that conversation is allowed to drift a little in the retelling. That one figure is not. The fix for the middle problem is about order, not content. Put key findings first in an aggregated input, not buried wherever they happened to land. Use explicit section headers so the model has structural cues, not just prose, guiding it to what matters. You are not fighting the lost in the middle effect by writing more clearly. You are fighting it by not putting anything you care about in the middle to begin with. Move the finding, not the sentence around it. The fix for field bloat is trimming at the source. Keep the five relevant fields from that order lookup. Drop the other thirty five before they ever accumulate in context, not after. Trimming after the fact means you already paid the token cost and already diluted attention with everything you were about to discard. Trim at the boundary, the moment the tool result comes back. The thirty five unused fields never get the chance to cost you anything at all. There is a basic requirement underneath all of this, easy to assume rather than actually verify. The complete conversation history has to go back in every subsequent request. The API does not remember anything between calls on its own. Drop part of that history by accident, trying to save tokens, and you have not trimmed context. You have broken the model’s ability to reason coherently about anything that came before the gap. Trimming and forgetting look identical from the outside, right up until the model needs the thing you dropped. One more piece belongs with task statement five point one, and it connects straight back to two episodes ago. Say an upstream agent hands off to a downstream one with a tight context budget. Send structured facts and citations. Not verbose reasoning chains. A downstream agent with limited room does not need to see how you arrived at a conclusion. It needs the conclusion, with enough metadata, dates, sources, to actually use it correctly. Send the reasoning anyway, and you have not helped the downstream agent. You have just spent its limited budget on your own thinking, instead of its own. Now task statement five point four. It is the same problem, at the scale of a whole investigation instead of one conversation. Long codebase exploration degrades in a specific, recognizable way. The agent starts giving answers about typical patterns instead of the actual class it found forty minutes ago. That is not the model getting worse. It is the model losing track of specifics it already established, and quietly substituting a generic default instead. A default sounds plausible. It sounds like it belongs. It is also not the actual class. Confidently generic is a worse failure than an honest I don’t remember would have been. At least an honest gap tells you to go look again. A confident, wrong default tells you nothing is wrong at all. Scratchpad files are the direct countermeasure. Have the agent write down key findings as it discovers them. Do not just hold them in a fading context window. Reference that file for later questions, instead of relying on memory of a conversation that keeps growing. A written fact does not degrade the way a summarized one does. Delegation helps here too, the same Explore subagent idea from the CI episode, applied to investigation instead of implementation. Spawn a subagent to answer one specific question. Find every test file. Trace the refund flow’s dependencies. The main agent holds only the high level picture. The verbose legwork happens somewhere it cannot dilute the coordination layer. The main agent never sees every file the subagent opened along the way. It sees the answer to the one question it actually asked. Between phases, summarize what the first phase actually found before spawning subagents for the next one. Inject that summary into their starting context. Each phase should start knowing what the last one learned, not starting cold and rediscovering it. Rediscovery is not thoroughness. It is the same work, paid for twice. A summary handed forward costs a few sentences. Rediscovering it later costs the whole exploration again. For anything long running enough to actually crash, structured state exports matter. Each agent writes its state to a known location. The coordinator, on resume, loads a manifest and injects it back into the relevant prompts. A crash then costs you the time since the last checkpoint, not the whole investigation from scratch. Six hours of exploration surviving a crash intact is the entire point of writing the manifest in the first place. Skip that step, and a crash at hour five is not a delay. It is the whole day, starting over. And when context is genuinely filling with exploration you no longer need, slash compact reduces it directly. It does not let an already strained context window get worse, turn after turn. Let me work a question. A coordinator running a six hour codebase migration keeps its full exploration history in context the entire time. Around hour four, it starts describing files it examin

    Episode 10. Context That Survives
  4. Sep 4

    Episode 9. Structured Output, Schemas, Validation, Batch

    Domain 4, task statements 4.3, 4.4 and 4.5. An extraction pipeline invents a clean, confident, entirely fake invoice number. Why a valid schema does not mean a true answer, nullable fields and enum escape hatches as fabrication controls, when a retry can and cannot fix a bad extraction, and when the Message Batches API saves money instead of costing you a blocked engineer. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 The invented invoice number * 0:59 Tool use with JSON schemas * 1:33 Syntax errors versus semantic errors * 2:23 Nullable fields against fabrication * 3:01 Enum escape hatches: unclear and other * 3:39 tool_choice any versus forced, for extraction * 4:15 Retry with the specific validation error * 4:40 Where retry stops working * 5:20 Self-checking schemas: totals and conflicts * 6:08 The detected-pattern field * 6:29 The Message Batches API * 6:54 Matching batch to a workload’s patience * 7:33 The multi-turn tool calling limit * 7:58 Sizing submission windows to an SLA * 8:30 custom_id and resubmitting failures * 8:52 Testing prompts on a sample first * 9:16 Worked question * 10:09 Why the other three options are there * 11:10 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 4.3, 4.4 and 4.5) * Structured outputs * Message Batches API Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The extractor pulled an invoice number off a document that never had one. Not garbled. Not close. A clean, confident, entirely invented number, sitting in the output field exactly where a real one belonged. The schema was valid. The JSON parsed perfectly. The number was fiction. Domain four continues. Task statements four point three, four point four and four point five today. How to actually guarantee structured output. What to do when it comes back wrong. When batch processing is the right tool, and when it is a trap. Start with the fix for syntax, because it is a real fix, and it is not the whole fix. Tool use with a JSON schema is the reliable route to structured output. Define the schema as a tool’s input, and the model returns data that already matches it. No more hand parsing a wall of prose, hoping the brackets close where you expect them to. That old approach failed in a very specific way. The model got the content right and the format wrong. A regex somewhere downstream choked on a comma it did not expect. Here is the boundary you need to hold in your head, because the exam draws it sharply. Strict schemas eliminate syntax errors. They do not eliminate semantic errors. A schema happily accepts a total that does not match the sum of its own line items. It accepts a value sitting in the wrong field entirely, formatted correctly, meaning nothing. Syntax and truth are different problems, and a schema only ever solves the first one. A well formed lie is still a lie. The schema just makes sure it is a well formed one. That gap is exactly where the invented invoice number lives. Nothing about a well formed schema stops a model from confidently filling a required field with something plausible. Not when the real value simply is not in the source document. The direct countermeasure: make fields optional, not required, whenever the source document might not contain the information. A nullable field can come back empty. A required field cannot. A model asked for something that must exist, and does not, will sometimes invent rather than admit absence. Optionality is not a weaker schema. It is the schema being honest about what a real document can and cannot promise. A required field is a promise the source document did not actually make. Do not force the model to keep a promise you made on the document’s behalf. Enums close a related gap. Real documents have messy middles, categories that almost fit. Forcing a fixed enum with no escape hatch pushes the model to pick the closest wrong answer, rather than say so. Add unclear as a legitimate value. Add an other option paired with a free text detail field. Now ambiguity has somewhere honest to go, instead of getting silently rounded off to whichever category happened to be nearby. A category with no honest exit does not remove ambiguity. It just hides it inside a wrong answer that looks exactly like a right one. Tool choice comes back into play here, and this time in the extraction context specifically. Set it to any when you are working across multiple possible extraction schemas. Use it when you do not yet know which one this document actually needs. That guarantees a tool gets called, without locking in which one ahead of time. Force a specific tool by name when a step genuinely has to run first. Extract metadata forced ahead of any enrichment step. Same ordering logic from a few episodes back, just applied here to extraction instead of research. Now task statement four point four, and this is where validation actually earns its keep or reveals it never had any. When a structured output fails validation, the fix is not a blind retry. It is a retry that includes the original document, the failed attempt, and the specific validation error. The model gets something concrete to correct, rather than a second blind guess. But know exactly where that retry stops working, because this is the single most testable idea in the task statement. A format error can be fixed on retry. A structural mismatch can be fixed on retry. Information that was never in the source document cannot be fixed on retry. No matter how many times you ask. There is nothing there to find. Recognizing which failure you are looking at, before you spend a second call on it, is the actual skill. Send a format error back for retry, and the model usually fixes it. Send an absence back for retry, and you have paid for a second identical guess. There is a self correcting schema pattern worth knowing by name. It catches the invented number problem at the design level, instead of after the fact. Extract a calculated total alongside a stated total, and flag it when they disagree. Add a conflict detected boolean for exactly this kind of internal contradiction. Two numbers that should match and do not is not noise. It is the document telling you something is wrong before a human ever has to notice by hand. The schema is not just capturing data anymore. It is capturing whether the document agrees with itself. That is precisely the signal that would have caught a fabricated invoice number. The moment it failed to reconcile with the line items sitting right next to it. One more habit worth building in from day one: a detected pattern field on every finding. When developers dismiss a flagged issue, that field lets you analyze which patterns keep generating false positives. Systematically. Instead of a vague sense that the tool is noisy, without ever being able to say exactly where. Last task statement: the Message Batches API. The trap the exam is watching for is not technical. It is a manager’s instinct. Batch cuts cost roughly in half. It runs within a processing window up to twenty four hours. It carries no guaranteed latency SLA, which is the detail that decides everything else in this task statement. Match the tool to the workload’s actual patience, not to the discount. An overnight report has all night. A weekly audit has all week. Nightly test generation has until morning. These tolerate the twenty four hour window without anyone noticing it happened. A pre-merge check has an engineer sitting there, blocked, waiting on the result before they can ship. Batch it, and you have traded a small cost saving. For an engineer staring at a spinner with no idea when it will resolve. Half price on a merge that never happens is not a discount. It is a stalled team. One mechanical limit worth knowing cold: batch does not support multi turn tool calling inside a single request. No mid-request tool execution and return. Say your workflow genuinely needs that back and forth. Batch is not an option for it, regardless of how patient the workload otherwise is. No amount of patience buys back a capability the API does not offer inside a batch request. Sizing submission windows against a real SLA is arithmetic, not guesswork. Promise a thirty hour turnaround. A twenty four hour processing window means you can only submit every four hours or so. That still lands inside your own promise, with room to spare. Do the math against the SLA you actually made, not the SLA you hope batch delivers on a good day. Batch has no obligation to move faster just because you are counting on it to. Custom ID fields are what make failure handling tractable at scale. Say some documents in a batch of a hundred fail. Resubmit only the failures, identified by their custom ID, with whatever fix the failure actually needs. Chunk an oversized document. Do not resubmit the entire hundred blind. And before any of that: refine your prompt against a small sample first. Test on ten documents before you commit to submitting ten thousand. A prompt that is seventy percent right at sample size is a warning, not a surprise you want at scale. Discover that after paying for the full batch, and you are resubmitting three thousand failures, one expensive cycle at a time. Let me work a question. A team extracts structured data from scanned invoices. On invoices missing a purchase order number, roughly one in six, the extractor retur

    Episode 9. Structured Output, Schemas, Validation, Batch
  5. Sep 4

    Episode 8. Precision Prompting, Criteria, Few Shot, Multi Pass

    Domain 4, task statements 4.1, 4.2 and 4.6. A review bot with a high false positive rate gets turned off by the team that built it. Why be conservative fails as an instruction, what makes a few shot example actually teach judgment instead of pattern matching, and why the same session that wrote code is the wrong session to review it. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 The review bot the team turned off * 1:01 Why be conservative fails * 1:44 Explicit categorical criteria * 2:38 Named categories over confidence thresholds * 3:51 Severity levels with code examples * 4:32 Few shot: examples are for ambiguity * 4:53 Reasoning versus answer keys * 5:29 Format consistency via examples * 5:50 Tool selection and messy documents * 6:52 Why self review is structurally biased * 7:40 The fix: an independent instance * 8:05 Multi pass: per-file and cross-file * 8:50 Confidence alongside findings * 9:16 Worked question * 10:22 Why the other three options are there * 11:09 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 4.1, 4.2 and 4.6) * Few shot prompting Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The team turned the review bot off. Not because it never caught anything real. Because for every real bug, it flagged four things that were not bugs at all. Engineers stopped reading its comments. A tool that is wrong most of the time trains people to ignore it, including the times it happens to be right. Domain four, prompt engineering and structured output, twenty percent of the exam, tied with domain three. Task statements four point one, four point two and four point six today. Precision, few shot examples, and how to structure a review so it does not fall apart at scale. Start with the fix somebody tried first, because it is the fix almost everyone tries first, and it does almost nothing. Be conservative. Only report high confidence findings. Those instructions sound like precision. They are not criteria. They are vibes with a confident tone. The model has no fixed definition of conservative to apply consistently. It applies its own shifting sense of it, turn to turn. Ask ten reviewers what conservative means and you get ten thresholds. Ask the model the same vague question ten times, across ten different files, and you get the same drift. What actually works is explicit, categorical criteria. Not check that comments are accurate. Flag a comment only when the claimed behavior contradicts what the code actually does. That is a test the model can run the same way every time. It names a specific condition, rather than a feeling. One names a condition. The other names a mood. Here is why this matters more than it sounds like it should. False positives do not stay contained to the category that produced them. A high false positive rate in one category burns trust in every category, including the ones the tool gets right constantly. Once someone stops reading the output, accuracy in the categories they never see does not help them at all. A correct finding nobody reads has the same practical value as no finding at all. Two moves fix this in practice. First, write specific criteria that state what to report and what to explicitly skip. Bugs and security issues, report. Minor style preferences, local team patterns, skip. Not confidence thresholds. Actual categories, named. A confidence threshold asks the model to rate its own certainty. Certainty is exactly the thing an overconfident model gets wrong most often. A named category asks it to check a fact instead. Second, and this is the one people resist, because it feels like giving something up. If one category is genuinely noisy, turn it off. Temporarily disable the worst offender while you fix its prompt, rather than let it keep eroding trust in the categories working fine. A team will forgive a tool for not covering everything. They will not forgive a tool that cries wolf on every single review. A gap is honest. Everyone knows the tool has not looked there yet. Coverage lost is a gap you can fill later. Trust lost is a much longer road back, and most teams never bother walking it. They just stop reading the bot. Severity classification gets the same treatment. Do not just say rate this critical, high, or low. Give each severity level a concrete code example of what belongs there. A shared example is worth more than a shared adjective. Two people, or two runs of the same model, read an adjective differently. They read an example the same way. Show what critical looks like in actual code, a hardcoded credential, an unbounded loop that can hang a request. Show what low looks like the same way, a missing comment, a variable name that could be clearer. The adjective invites debate. The example ends it. Now few shot prompting, task statement four point two, and here is the core idea worth holding onto: examples are for ambiguity. Say detailed instructions alone still produce inconsistent output. That is the signal you need two to four targeted examples, not a longer paragraph of more instructions. The examples that actually help are not just answer keys. They show reasoning. An example that demonstrates why one tool was chosen over a plausible alternative teaches the model something durable. It generalizes that judgment to a case it has never seen. An example that just shows an answer with no reasoning attached teaches pattern matching to that one case, and nothing else. Bare answers train a lookup table. Reasoning trains a judgment. One scales to new cases. The other only ever repeats the old ones. Format consistency is the other place few shot earns its keep. Show two or three examples in the exact output shape you want. Location, issue, severity, suggested fix. The model locks onto that shape far more reliably than a written description of the same fields ever manages. Few shot also earns its place wherever a request is genuinely ambiguous rather than just varied. Tool selection between two plausible tools is one such case. Show two or three examples of a borderline request. Pair each with the tool that was actually chosen, and a line of reasoning for why. The model is not memorizing those exact requests. It is learning the shape of the decision. That is what lets it handle the next borderline case correctly, the one none of your examples ever showed it. There is a sharper use case worth knowing by name: extraction from messy documents. Vary the examples across the kinds of structural variety you actually expect. Inline citations against full bibliographies. A methodology section against details buried in a table. Do that, and empty or null extractions on the fields that matter drop hard. The model has seen the shape of the mess before, not just the shape of a clean case. Last task statement, and it belongs here because it is really about trust in a different form: the review itself. Task statement four point six, multi instance and multi pass architectures. Here is the limitation the exam wants you to know cold. A model that just generated some code keeps the reasoning it used to write that code, in the same session. Ask it to review its own work, and it is reviewing decisions it already committed to. With the confidence of having just made them. Telling it to think harder, or extending its reasoning time, does not remove that bias. It is still the same session, holding the same committed context. Nothing about thinking longer changes whose decisions are under review. The fix is not a smarter self review prompt. It is a second, independent instance, one with no memory of writing the code, looking at it fresh. That independent instance catches things the generator does not. Not because it is smarter. Because it never had a reason to defend the first draft’s decisions. It walks in cold. Nothing to defend, nothing already decided. Multi pass review solves a different problem, the one from episode one, back again in this domain. Split a large review into a per file pass for local issues. Add a separate cross file pass for how the pieces interact. Do both in one sweep across everything, and attention dilutes exactly the way it did with the fourteen file pull request. Split it, and each pass actually gets to focus. One pass reads a single file closely. A second pass reads across files for exactly the kind of contradiction a narrow, file by file view would never surface. Neither pass is doing the other’s job. That is exactly why both together outperform one pass trying to do everything at once. One more piece worth naming: confidence alongside findings. Have the review self report a confidence level on each individual finding, not just an overall verdict. That number is what lets you route intelligently later. Low confidence findings go to a human. High confidence ones go straight through. Nobody treats every finding the same, regardless of how sure the model actually was. Let me work a question. A code review tool flags issues with instructions telling it to be thorough and to use good judgment about what matters. Two different files with nearly identical patterns get inconsistent verdicts, flagged in one, ignored in the other. What is the most direct fix. Option one, tell the model to be more consistent across files. Option two, replace the judgment based instruction with explicit criteria. Name which patterns to flag and whi

    Episode 8. Precision Prompting, Criteria, Few Shot, Multi Pass
  6. Sep 4

    Episode 7. Plan Mode, Iteration, and CI

    Domain 3, task statements 3.4, 3.5 and 3.6. A CI pipeline hangs for forty minutes waiting for a keystroke that will never come. When plan mode earns its cost and when it is just ceremony, four techniques for iterating well, and the exact flags that make Claude Code behave in an unattended pipeline. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 The pipeline that hung for forty minutes * 0:59 Plan mode: architectural weight * 1:41 Direct execution: the other legitimate choice * 2:23 The trap: size versus approaches * 2:42 Combining plan mode and direct execution * 3:13 The Explore subagent * 4:14 Four iteration techniques * 4:23 Concrete input and output examples * 4:49 Test-driven iteration * 5:29 The interview pattern * 6:02 Batching interacting fixes versus sequential * 7:03 The -p flag * 7:38 Structured CI output: json and json-schema * 8:09 CLAUDE.md as CI context * 8:36 Why an independent instance reviews better * 9:18 Giving a re-run memory of prior findings * 10:00 Worked question * 11:12 Why the other three options are there * 11:50 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 3.4, 3.5 and 3.6) * CLI reference (-p, --output-format, --json-schema) * Headless and CI use Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. A pipeline sat hung for forty minutes. Nothing crashed. Nothing timed out on its own. Claude Code was simply sitting there, waiting politely for someone to type an answer to a question. Nobody was watching the machine. The job was supposed to run unattended overnight, and finish before anyone woke up. Domain three continues. Task statements three point four, three point five and three point six. Today: when to plan before you build. How to iterate well once you are building. What it actually takes to run Claude Code somewhere nobody is sitting at a keyboard. Start with plan mode. The exam tests a judgment call here, not a mechanism. Plan mode is for tasks with real architectural weight. Large scale changes. More than one valid approach on the table. Decisions that touch the shape of the system, not one function inside it. A migration touching forty five files. A choice between two infrastructure approaches with genuinely different tradeoffs. A restructuring that touches how a dozen services talk to each other. Commit to code before thinking the approach through, and you risk expensive rework. Plan mode lets that thinking happen safely, before a single line changes. Direct execution is the other half, and it is just as legitimate. Not a lesser choice. A bug fix confined to one file, where a stack trace already points at the exact line. Adding one validation check to one function. Well scoped, well understood changes. A planning phase here adds ceremony, not safety, because there was never more than one reasonable way to make the fix. Size alone will fool you here. A one line change to a shared authentication check can carry more architectural weight than a five hundred line refactor confined to one module. Count the approaches on the table, not the lines in the diff. The exam’s favorite trap is not asking you to define plan mode. It describes a task and asks whether it warrants planning at all. The tell is not the size of the diff. It is whether multiple valid approaches exist, or the path was already obvious the moment the bug was found. Real work often needs both halves, not one or the other. Plan mode for the investigation. Direct execution once the plan is settled. Work out how a library migration should actually proceed, which files it touches and in what order. Once that plan is solid, switch to direct execution and carry it out. Planning and executing are not phases you pick once and stay in. They are two tools you hand off between, inside the same piece of work. One more piece belongs here, and it solves a problem specific to investigation: context getting eaten by discovery before you have even started building. A multi phase task often needs a verbose exploration phase first. Reading a lot of files. Tracing a lot of imports. Do all of that directly in your main conversation, and your context window fills up with exploration before implementation has even started. The Explore subagent exists for exactly this. It isolates the discovery work. Only a summary comes back. Picture mapping an unfamiliar service across thirty files before touching any of them. Do that through the Explore subagent, and your main conversation gets a clean summary of what the service does and how its pieces connect. Not thirty files worth of raw exploration you will never need to see again. The context you saved is context still available for the part of the job that actually matters: writing the change itself. Now task statement three point five. Getting better output once you are actually iterating, rather than deciding whether to plan. First technique: concrete input and output examples. A prose description gets interpreted inconsistently across attempts more often than people expect. Two or three actual examples, input and correct output, communicate a transformation far more reliably than another paragraph trying to nail it down. Show the shape. Do not just describe it. Second technique: test driven iteration. Write the test suite first. Expected behavior, known edge cases, performance requirements that matter. Then iterate by sharing the actual test failures back, rather than describing the problem in your own words each time. A failing test is a precise signal. Your own summary of the problem often is not. Hand the model a stack trace and a failing assertion, and it knows exactly what wrong looks like. Hand it a paragraph of your own words, and it is reconstructing your intent from a description, one step removed from the actual failure. Third technique, and the one people underuse the most: the interview pattern. Before implementing something in an unfamiliar domain, have Claude ask you questions first, rather than diving straight into code. Cache invalidation strategy. Failure mode handling. Edge cases around concurrent access. A developer moving fast might never think to specify any of it up front. Good clarifying questions surface it before it gets baked into a wrong design, instead of after. Fourth, and this is the one with the sharpest exam signature: batching fixes versus fixing sequentially. Issues that interact, where changing one thing affects how another part of the fix should work, belong in one detailed message together. Fix them one at a time, and each fix risks undoing or complicating the next. Issues that are genuinely independent cost nothing to handle sequentially. The exam question here is rarely about the technique. It is about judging correctly whether the issues in front of you actually interact. Get that judgment wrong in either direction and it costs you. Batch independent issues together, and you have written one long, tangled message where a series of short ones would have been clearer and easier to verify one at a time. Split interacting issues apart, and the second fix quietly undoes something the first one depended on, because nobody told it the two were connected. Last task statement: getting Claude Code into a continuous integration pipeline, which is exactly where the forty minute hang I opened with happened. The fix has a name, and it is the single most testable fact in this task statement. The dash p flag, sometimes written dash dash print, runs Claude Code non interactively. No prompt ever waits for a keystroke, because there is no keystroke coming in an automated pipeline, and the flag exists specifically so the tool never tries to ask for one. Two more flags make the output usable by other software, not just readable by a human. Output format json turns the response into something your pipeline can actually parse, instead of prose you would have to scrape. Json schema goes further, enforcing that the structure matches a schema you define. That is what lets you post specific findings as inline pull request comments, instead of one giant unstructured comment block. CLAUDE.md does real work here too. It is not just for interactive sessions. In a CI context, it is how a headless invocation learns your testing standards, your fixture conventions, your review criteria, the same way it would teach any interactive session. A CI run with no CLAUDE.md context is reviewing code against no stated standards beyond generic good practice. There is a subtler point here, and it changes how you would actually architect a review pipeline. The same session that wrote a piece of code is measurably worse at reviewing it than an independent session would be. It already committed to its own reasoning while writing it. Reviewing your own work means re-examining decisions you already made confidently, and that is harder than looking at someone else’s code fresh. The consequence: your CI review step should be a separate, independent invocation, never a continuation of whatever session generated the change. Different eyes, every time, even when those eyes belong to the same model. Two more habits round this out. When a review job runs again after new commits, include the prior findings in context, and tell it to flag only what is new or still unresolved. Skip that, and every re-run repeats comments a human already read and fixed or dismissed. Do tha

    Episode 7. Plan Mode, Iteration, and CI
  7. Sep 4

    Episode 6. Configuring Claude Code, Memory, Rules, Commands, Skills

    Domain 3, task statements 3.1, 3.2 and 3.3. A new hire’s Claude Code ignores half the team’s conventions, and the reason is not code, it is configuration scope. The three CLAUDE.md levels, splitting a sprawling file with @import and a rules directory, commands versus skills, and why a scattered convention needs a glob-scoped rule instead of a directory-level file. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:24 The new hire’s broken conventions * 0:44 Three configuration levels * 1:36 Why user-level settings never travel * 2:08 The fix: check which level it lives at * 2:33 The /memory command as diagnosis * 3:19 Keeping CLAUDE.md from sprawling: @import * 4:09 The rules directory, split by topic * 4:33 Commands: project versus personal scope * 4:59 Skills and their frontmatter * 6:04 Personal skill variants * 6:45 Skill or CLAUDE.md: the real judgment call * 8:13 Path-scoped rules and glob patterns * 9:50 Worked question * 10:55 Why the other three options are there * 11:56 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 3.1, 3.2 and 3.3) * CLAUDE.md hierarchy and /memory * Skills and SKILL.md frontmatter Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. A new engineer joins the team. Opens Claude Code on a repository everyone else has worked in for a year. Gets code that ignores half the team’s conventions. Claude Code is not broken. The conventions live in a file on someone else’s machine. Nobody ever put them anywhere this engineer’s copy could see. Domain three, twenty percent of the exam, tied with domain four. Task statements three point one, three point two and three point three. Where configuration lives. How commands and skills get built. How a rule finds the files it actually applies to. Start with the hierarchy, because the new hire’s problem is a hierarchy problem. Three levels. User level, in your own home directory, applies only to you. Project level, in the repository itself, applies to everyone who checks it out. Directory level, inside a specific subdirectory, applies to work happening there. A CLAUDE.md sitting inside a payments folder shapes how Claude Code behaves while working inside that folder, and nowhere else. Move to a different part of the codebase, and that file has no say. Here is the fact that explains the whole broken experience, and it catches experienced engineers too, not just newcomers. User level settings are personal. They do not travel through version control. A senior engineer can spend six months tuning their own setup and never write a line of it into the project file. All six months of tuning belongs to one person, on one machine. A new teammate inherits none of it. Not their fault. There was never anything in the repository to inherit. The fix is not a lecture about onboarding. It is a habit. When Claude Code ignores a convention you know exists, check which level it actually lives at. Sitting in someone’s personal file, it was never going to travel with the repository. That five second check saves the hour someone would otherwise spend rewriting a convention that already existed, just in the wrong place. Claude Code gives you a tool for exactly this. The slash memory command shows which memory files actually loaded this session. Run it, and you see the real list. User level file, present or absent. Project level file, present or absent. Every rule file that actually fired. Say two engineers report different behavior from the same prompt. Both think they are running the same setup. Run slash memory in each of their sessions and compare. Nine times out of ten, one of them has a personal file the other does not, and the mystery is solved in under a minute. Inconsistent behavior across sessions or teammates, that command tells you what is really in effect, not what you assume. Now, keeping a large CLAUDE.md from turning into a monster. Two mechanisms. At import lets one CLAUDE.md pull in another file. A payments package imports the standards on money handling and audit logging. A frontend package imports the standards on components and accessibility. Neither maintainer reads the other’s rules. The shared file gets written once. The alternative to at import is worse in a specific way worth naming. Paste the same standards into every package’s CLAUDE.md by hand, and the moment one standard changes, you are hunting down every copy to update it. At import means the source of truth lives in one file, and every package that references it gets the update automatically, the next time it loads. The other mechanism is a rules directory. Instead of one giant file, split conventions by topic. Testing in one file, A P I conventions in another, deployment in a third. Each file stays short enough that someone can actually find the rule they need. Nobody has to scroll past deployment conventions to find the one testing rule they actually came looking for. Task statement three point two. Commands and skills. Easy to blur if you have not built both. Commands live in a commands directory. Project scoped, shared through version control, for the whole team. Personal scoped, in your own home directory, just for you. Write a project scoped command once, and it is available to everyone who checks out the repository. Skills carry their own file, SKILL dot M D, with frontmatter worth knowing by name. Context fork comes first. A skill that produces a lot of noise, a full codebase analysis, ten brainstormed approaches, does not have to dump that into your main conversation. Run it with context fork instead. The work happens in its own sub-agent context. Only the result comes back. Allowed tools restricts what a skill can touch while it runs. A skill whose whole job is writing a report gets scoped to file writes only. Nothing else. A bug in that skill cannot reach further than the report. And argument hint prompts a developer for whatever parameter the skill needs, the moment they invoke it without one. Invoke a deploy skill with no target environment named, and instead of guessing, or failing silently, it asks. Staging or production. One clear question, instead of a skill that ran against the wrong environment because nobody told it which one. One more pattern. Want your own version of a shared skill, without touching the one everyone else uses. Give it a different name in your personal skills directory. Your version sits alongside the team’s, not in place of it. Say the team’s release skill always runs a full test suite before tagging a version. You want a faster personal variant for your own local drafts, one that skips the slow integration tests. Name it something else. Release dash draft, say. Now it lives next to the team’s release skill, not instead of it. Nobody else’s workflow changes. Nobody has to know you built a shortcut for yourself. The real judgment call in this task statement: skill, or CLAUDE.md. CLAUDE.md loads every session, automatically. A skill loads on demand. Universal standards, the ones that should shape every interaction, belong in CLAUDE.md. A specialized workflow, one only a fraction of sessions ever need, belongs in a skill. Think of it as the difference between a habit and a tool in a drawer. A habit runs every time, without being asked. A tool in a drawer only costs you anything the moment you reach for it. Load it into every session anyway, and you are paying for relevance nobody asked for. Commit message formatting goes in CLAUDE.md, since every session might eventually commit. A legacy API migration goes in a skill, since only the sessions doing that migration will ever need it. Get the judgment call backwards and both directions hurt. Put a narrow, rarely used workflow into CLAUDE.md, and every single session pays a small tax for something almost nobody needs that day. Put a universal standard into a skill instead, and it silently stops applying the moment nobody remembers to invoke it. Neither failure looks dramatic. Both just quietly cost you, one token at a time, or one missed convention at a time. Last task statement, and it has the cleanest exam signature of the three. Path specific rules. A rules file carries Y A M L frontmatter with a paths field, glob patterns, and the rule activates only when you are editing a matching file. This beats a subdirectory CLAUDE.md in one specific way, and the exam loves testing files to make the point. Say your tests are scattered across a dozen directories, not gathered into one folder. A subdirectory CLAUDE.md cannot reach all of them. There is no single directory to put it in. A glob scoped rule can. One pattern, matching every test file by name, loads the convention wherever that file happens to live. None of this makes directory level CLAUDE.md the wrong tool everywhere. A folder that genuinely contains one coherent thing, a specific microservice, a specific package, is exactly what it is built for. The failure only shows up when the convention and the folder structure do not line up. When what you are trying to govern is a file type scattered across the tree, rather than a place in it. There is a second reason this matters beyond correctness: tokens. A rule that only loads for matching files means the model is not carrying testing conventions while it works on something that has nothing to do with tests. Working on a payments function, editing zero test f

    Episode 6. Configuring Claude Code, Memory, Rules, Commands, Skills
  8. Sep 4

    Episode 5. MCP in Production, Errors, Scoping, Built Ins

    Domain 2, task statements 2.2, 2.4 and 2.5. An agent stuck retrying a permission error against a tool that only ever says operation failed. The four MCP error classes, project versus user scoped server configuration, MCP resources as content catalogs, and picking the right built-in tool for the job instead of the familiar one. Independent and unofficial. This series is not affiliated with, sponsored by, or endorsed by Anthropic. Nothing in it is exam content. Every practice question was written for this show against the published exam guide, which is a free public document linked below. Chapters * 0:00 Cold open and disclaimer * 0:23 An agent retrying a locked door * 0:52 The isError flag, and its limit * 1:37 Four error classes * 2:17 Structured metadata: category, retryable, description * 3:18 Local recovery versus propagating upward * 4:11 Access failure versus a valid empty result * 5:07 Project versus user scoped MCP config * 6:02 Credentials via environment variable expansion * 6:33 Writing MCP descriptions the agent will prefer * 7:13 Community servers over custom ones * 7:37 MCP resources as content catalogs * 8:07 Grep, Glob, Read, Write, Edit * 8:43 The Read plus Write fallback * 9:16 Exploring an unfamiliar codebase * 9:57 Worked question * 11:00 Why the other three options are there * 11:51 What the exam will ask, and next time Sources * Claude Certified Architect, Foundations Exam Guide, Version 1.0 (section 6, task statements 2.2, 2.4 and 2.5) * MCP in Claude Code Transcript This is Passing C C A R F, an independent study companion for the Claude Certified Architect, Foundations exam. Sixty items, a hundred and twenty minutes, and a published blueprint that tells you almost exactly what it is going to ask. Independent and unofficial. Not affiliated with or endorsed by Anthropic. No exam content. The agent hit a permission error on a backend call, and it tried again. Then it tried again. Four more times, each with a slightly different approach, each one failing for the exact same reason, because the tool’s response, every time, was two words: operation failed. Nothing about why. Nothing about whether trying again could possibly help. The agent had no way to know it was banging on a locked door, so it kept knocking. Domain two continues. Task statements two point two, two point four and two point five. Error responses, server scoping, and the built in tools everyone already has and half the time reaches for wrong. Start with the error, because it is the cleanest version of a pattern you have already seen in this series: a uniform response hides information the agent needs to act well. The Model Context Protocol has a mechanism for this, an is error flag that marks a tool result as a failure rather than a success. That flag alone tells the model something went wrong. It does not tell the model what kind of wrong, and that distinction is the entire episode. There are four classes of error worth telling apart, and the exam wants you fluent in all four. Transient errors, a timeout, a service that is briefly unavailable, where trying again in a moment might genuinely work. Validation errors, the input itself was malformed, where retrying the identical call will fail identically forever. Business errors, the operation was understood perfectly and refused on purpose, a policy violation, not a glitch. And permission errors, which is exactly what opened this episode, where the caller is not allowed to do this at all, and no number of retries changes that. Notice what those four categories buy you that a flat operation failed cannot. They tell the agent whether retrying is even worth attempting. A generic failure response leaves that question unanswered. So the agent either gives up too early on something recoverable, or burns time retrying something that was never going to succeed no matter how many times it asked. The fix is structured metadata, and the shape of it is worth knowing by name. An error category field, transient, validation, business or permission. An is retryable boolean, so the agent does not have to guess from the category alone. And a human readable description that can be surfaced to an actual customer when the failure is the kind of thing a customer needs to hear about. For business errors specifically, that readable explanation matters even more. A policy violation is not a bug to route around. It is information the person on the other end deserves to receive in plain language. Here is where the four category, is retryable pattern actually saves you real cost, and it is the same logic as the enforcement episode a few back. Local recovery belongs inside the subagent that hit the failure. If a call times out and the category is transient, retry it there, quietly, without ever bothering the coordinator. What should travel upward to the coordinator is only the failure that could not be resolved locally. It should travel with two things attached: what was actually attempted, and whatever partial results already exist. A coordinator that receives a bare failure notice with no partial results has to start that piece of the investigation from nothing. A coordinator that receives a structured failure with partial results attached can often keep going with what it already has. One more distinction inside this same territory, and it is subtle enough that people genuinely mix it up under time pressure. An access failure and a valid empty result are not the same thing, even though both can look, from a distance, like nothing came back. An access failure means the query could not run, permission denied, service unreachable, something is actually broken. A valid empty result means the query ran perfectly and there is nothing in your data matching that particular ask. Treating an empty result as a failure means retrying a search that was never going to return anything different. Treating a real failure as an empty result means quietly reporting success on a query that never actually executed. Both are wrong in opposite directions, and the fix for both is the same: know which one you are looking at before you decide what to do next. Now scoping, task statement two point four, and this is mostly about where configuration lives rather than what it says. Two levels. Project scoped configuration, in dot m c p dot Jason, is for shared team tooling, the servers everyone on the project needs and that travel with the repository in version control. User scoped configuration, in your home directory’s claude dot Jason, is for personal or experimental servers. Things you are trying out, that have no business being forced on every other person who checks out the repo. Both levels are active at once. Tools from every configured server, project and user, are discovered when the connection is made and are all available to the agent simultaneously. This is not an either or choice, it is two pools that both feed the same agent. Credentials belong in the project file without ever actually being in the project file, and the mechanism is environment variable expansion. Write a token as a reference to an environment variable rather than the literal secret, and the actual value is pulled from the environment at run time. This is how a team shares MCP server configuration in git without also sharing every team member’s API keys in git, which would otherwise be the obvious and disastrous alternative. There is a quieter skill buried in task statement two point four that is easy to skip past: writing MCP tool descriptions well enough that the agent actually prefers them over a built in. Say you connect a capable MCP tool for ticket lookups, and its description is thin. The agent may keep reaching for a generic built in search instead. The built in tool is the path of least resistance, and the MCP tool never explained why it is the better choice. The fix is the same one from two episodes back: write the description like it is the interface, because it is. And a related judgment call: reach for an existing community MCP server before building your own, for anything standard. Most teams do not need a custom Jira integration when a maintained community one already does the job. Save custom servers for the workflows that are genuinely specific to your team, where nothing off the shelf fits. Last idea in this task statement, and it is a nice one: resources. A resource is not a tool you call to take an action. It is a catalog you can browse, exposed by the server itself: a list of issue summaries, a documentation hierarchy, a database schema. Giving an agent access to that catalog directly means it does not need a round of exploratory tool calls first. It can see what exists before it decides what to actually ask for. Finally, task statement two point five, the built in tools you already have without connecting anything. Grep searches file contents, the tool for finding a function name, an error message, or every place a particular string appears across a codebase. Glob matches file paths by pattern, the tool for finding files by name or extension rather than by what is written inside them. Read and Write handle whole files. Edit makes a targeted change by matching unique surrounding text, and it is precise exactly because it insists on uniqueness. That insistence is also where Edit sometimes fails, and the fallback matters because it comes up constantly in real work. If the text you are trying to anchor an edit to appears more than once in a file, Edit cannot tell which occurrence you mean, and it will refuse rather than guess wrong. The reliable fallback in that situation is Read the whole file, make the change yourself in memory, then Write the full file back. Not a workaround, a documented fallback for exactly this case. There is also a way of working through an unfamiliar codebase that the exam rewards over the more obvious approach. Do not start by reading every file. Start with Grep to find entry points, the handful of

    Episode 5. MCP in Production, Errors, Scoping, Built Ins

About

An independent, unofficial study companion for the Claude Certified Architect, Foundations exam. Not affiliated with, sponsored by, or endorsed by Anthropic. brodynetworks.substack.com