← Back

20 Output Evaluation & Validation Practice Questions & Answers

Every Output Evaluation & Validation practice question from the Claude Certified Associate – Foundations Practice Test, with the correct answer and a short explanation.

Start practice test
  1. 1. A compliance analyst asks Claude to summarize a federal regulation, and the response includes citations to specific regulation sections. What should the analyst do before relying on those citations?

    • A.Trust them because they are formatted like real legal citations
    • B.Ask Claude in a follow-up message whether the citations are accurate
    • C.Verify only the citations that look unusual or unfamiliar
    • D.Look up every cited section in the official regulation text to confirm it exists and supports the claimAnswer

    Language models can fabricate citations that are perfectly formatted yet point to sections that do not exist or say something different, so formatting is no evidence of accuracy. Because Claude cannot reliably audit its own prior output, the only dependable check is to confirm each citation against the authoritative regulation text itself.

    Source: Anthropic, "Reduce hallucinations" (docs.claude.com/en/docs/build-with-claude); CCAO-F Exam Guide, Domain 2 (verifying citations against authoritative sources)Report a problem with this question

  2. 2. Which statement about hallucinations in Claude's outputs is accurate?

    • A.Hallucinated content is usually hedged with phrases like "I'm not sure", making it easy to spot
    • B.Hallucinations only occur in very long responses
    • C.Hallucinations are often fluent and confidently phrased, so tone is not a reliable signal of accuracyAnswer
    • D.Hallucinations are limited to numerical calculations

    A model generates fabricated details using the same fluent, assertive style it uses for correct facts, because it produces plausible text rather than retrieving verified records. That is why evaluation guidance says to verify substance against sources instead of judging reliability by how confident the writing sounds.

    Source: Anthropic, "Reduce hallucinations" — hallucinations can be stated with high confidence (docs.claude.com/en/docs/build-with-claude); CCAO-F Exam Guide, Domain 2Report a problem with this question

  3. 3. Claude drafts a summary of new industry regulations that will be circulated to a company's compliance team. What is the most appropriate step before sharing it?

    • A.Share it immediately since the compliance team will catch any errors later
    • B.Check the summary against the official regulation text and have a qualified compliance professional review itAnswer
    • C.Ask Claude to regenerate the summary and share whichever version is longer
    • D.Run the summary through a grammar checker before distribution

    Regulatory content is high-stakes: an inaccurate summary could lead the organization to violate a rule, so the verification standard must match the risk. Checking against the authoritative source catches factual drift, and expert human review is required because the output will inform compliance decisions Claude cannot be accountable for.

    Source: CCAO-F Exam Guide, Domain 2 (human review for high-stakes outputs); support.claude.com — verify important information from Claude before acting on itReport a problem with this question

  4. 4. Claude analyzed a sales spreadsheet and produced summary figures for an executive presentation. What is the best way to validate the analysis before the presentation?

    • A.Ask Claude to rate its own confidence in each number
    • B.Accept the figures if they roughly match your intuition about the business
    • C.Compare the figures to last year's presentation
    • D.Recompute or spot-check the key figures directly against the original dataAnswer

    The original dataset is the ground truth for any derived figure, so recomputing or spot-checking against it is the only method that can actually detect a calculation or extraction error. Intuition, self-reported confidence, and prior presentations can all agree with a wrong number, because none of them independently verifies the arithmetic.

    Source: CCAO-F Exam Guide, Domain 2 (fact-checking outputs against source data before business use)Report a problem with this question

  5. 5. Which scenario most clearly requires human review before the Claude output is used?

    • A.A casual internal chat message summarizing a meeting you attended
    • B.A customer-facing notice describing the legal terms of a refund policyAnswer
    • C.A first draft outline that the author will heavily rewrite anyway
    • D.A brainstorm list of ideas for an internal team offsite

    The need for human review scales with the consequences of an error: customer-facing statements about legal terms create binding expectations and legal exposure if wrong, so they must be verified by a person before release. The other outputs are low-stakes or already subject to the author's own revision, where an error is easily caught or harmless.

    Source: CCAO-F Exam Guide, Domain 2 (risk-based criteria for when human review is required); support.claude.com usage guidance on high-stakes contentReport a problem with this question

  6. 6. When reviewing a Claude-drafted job description for bias, what should the reviewer primarily look for?

    • A.Whether the vocabulary level is appropriate for the industry
    • B.Whether the description mentions the company's name often enough
    • C.Whether the language, requirements, or examples systematically favor or discourage particular groups of candidatesAnswer
    • D.Whether the text is shorter than descriptions written by humans

    Bias in generated content means systematic patterns that advantage or disadvantage groups — for example gendered wording, unnecessary requirements that screen out certain candidates, or one-sided framing — because the model reproduces patterns from its training data. Length, vocabulary level, and branding are style concerns, not bias, so a bias review specifically examines whether the content treats groups unevenly.

    Source: CCAO-F Exam Guide, Domain 2 (identifying bias in outputs); support.claude.com — Claude can reflect biases present in training dataReport a problem with this question

  7. 7. A Claude-generated report states a customer churn rate of 12% in the executive summary but 21% in the detailed findings section. What does this indicate and what should you do?

    • A.Average the two values and report 16.5%
    • B.It is an internal inconsistency; check the underlying data to determine the correct figure before using either numberAnswer
    • C.The detailed section is always more reliable, so use 21% without further checking
    • D.It is normal rounding variation; use the executive summary figure

    Two contradictory values for the same metric mean at least one is wrong, and neither the summary's position nor the detail section's length gives it authority — the source data does. Internal inconsistency is a classic quality signal that requires resolution against the ground truth, because guessing, averaging, or defaulting to one section can propagate a wrong number into decisions.

    Source: CCAO-F Exam Guide, Domain 2 (detecting internal inconsistencies and resolving against source data)Report a problem with this question

  8. 8. You need to fact-check a Claude output describing a government agency's filing requirement for businesses. Which source is most authoritative?

    • A.A discussion thread where practitioners share their experiences
    • B.The agency's own official website or the published regulation textAnswer
    • C.A news article that mentions the requirement in passing
    • D.A consulting firm's blog post summarizing the requirement

    The agency that issues and enforces a requirement is the primary source: its official publications define the rule rather than interpreting it secondhand. Blogs, forums, and news coverage are secondary sources that may be outdated, simplified, or simply wrong, so fact-checking guidance directs verification to primary, official sources whenever they exist.

    Source: CCAO-F Exam Guide, Domain 2 (fact-checking against authoritative/primary sources)Report a problem with this question

  9. 9. Claude produced a technically accurate but jargon-heavy analysis. You need to present it to senior executives. What is the best way to adapt it?

    • A.Add more technical detail to demonstrate the analysis is rigorous
    • B.Delete all numbers so the executives are not overwhelmed
    • C.Rework it to lead with conclusions and business implications, replacing jargon with plain language while preserving the factsAnswer
    • D.Present it unchanged so no technical nuance is lost

    Adapting output for an audience means changing presentation — order, vocabulary, and emphasis — without changing the underlying facts. Executives make decisions from conclusions and implications, so leading with those and removing jargon serves the audience, whereas leaving it unchanged, stripping the evidence, or adding complexity all fail either the audience or the accuracy requirement.

    Source: CCAO-F Exam Guide, Domain 2 (editing and adapting outputs for the intended audience)Report a problem with this question

  10. 10. In the Claude app, when is an artifact the better choice than an inline chat response?

    • A.When you want the response to be permanently unmodifiable
    • B.When the output is substantial, self-contained content the user will edit, reuse, or share, such as a document or code fileAnswer
    • C.When the answer is a short factual reply to a quick question
    • D.Whenever the response contains any numbers at all

    Artifacts exist to hold substantial, standalone content in a dedicated window where it can be iterated on, versioned, and shared, separating the deliverable from the conversation. Short conversational answers belong inline because moving them to an artifact adds friction without benefit, and artifacts are editable — not frozen — so permanence is not their purpose.

    Source: support.claude.com — "What are Artifacts and how do I use them?"; CCAO-F Exam Guide, Domain 2 (choosing output formats)Report a problem with this question

  11. 11. A team wants Claude's output to be ingested automatically by another software system. Which output format choice is most appropriate?

    • A.A bulleted summary of the highlights only
    • B.Structured data such as JSON that follows a defined schemaAnswer
    • C.A conversational answer with the data mentioned in sentences
    • D.Flowing narrative prose with rich formatting

    Machines parse data reliably only when it arrives in a predictable structure, so specifying a schema-conformant format like JSON lets the downstream system extract fields deterministically. Prose and bullet summaries force fragile text parsing and can silently drop required fields, which is why format choice should follow how the output will be consumed.

    Source: Anthropic docs — "Increase output consistency (JSON mode / structured outputs)" (docs.claude.com/en/docs/build-with-claude); CCAO-F Exam Guide, Domain 2Report a problem with this question

  12. 12. What does evaluating a Claude output for completeness primarily involve?

    • A.Confirming the response is at least several paragraphs long
    • B.Checking that every part of the request was addressed and that material caveats or limitations were not omittedAnswer
    • C.Verifying the response uses correct grammar throughout
    • D.Ensuring the response repeats the original question before answering

    Completeness is about coverage, not length: an output can be long yet skip half the request, or omit a caveat that changes how the answer should be used. Systematically mapping the response back to each element of the request — including necessary qualifications — is what catches silent omissions, which are a distinct failure mode from inaccuracy.

    Source: CCAO-F Exam Guide, Domain 2 (evaluating outputs for accuracy and completeness)Report a problem with this question

  13. 13. According to Anthropic's guidance, which prompting technique helps reduce hallucinations in Claude's responses?

    • A.Telling Claude to sound more confident
    • B.Explicitly giving Claude permission to say "I don't know" when it is unsureAnswer
    • C.Asking for longer, more detailed responses
    • D.Instructing Claude to always provide an answer even when uncertain

    Anthropic's hallucination-reduction guidance recommends explicitly allowing an "I don't know" response because a model pushed to always answer will fill gaps with plausible fabrications. Giving an uncertainty escape hatch changes the incentive: declining becomes an acceptable output, so the model is less likely to invent details to satisfy the request.

    Source: Anthropic, "Reduce hallucinations" — allow Claude to say "I don't know" (docs.claude.com/en/docs/build-with-claude)Report a problem with this question

  14. 14. A marketing team proposes publishing Claude-generated product descriptions directly to the company website without review. What is the strongest objection?

    • A.Claude's writing style is too informal for any website
    • B.Customer-facing factual claims carry legal and brand risk, so a human must verify accuracy before publicationAnswer
    • C.Generated text always contains grammatical errors
    • D.Publishing AI-generated text is not technically possible

    Product descriptions make factual claims — features, capabilities, compatibility — and a hallucinated claim published to customers can trigger false-advertising liability and brand damage. Because the model can produce confident errors, customer-facing publication is a high-stakes use where human verification is the control that must precede release; the other objections are factually wrong generalizations.

    Source: CCAO-F Exam Guide, Domain 2 (human review before customer-facing publication); support.claude.com — verify Claude's outputs for important usesReport a problem with this question

  15. 15. Claude describes a rule and notes it was "recently updated." Why should this specific claim receive extra scrutiny?

    • A.Because Claude never has information about any rules
    • B.Because recent updates are always minor and can be ignored
    • C.Because rules are updated too rarely for the claim to be plausible
    • D.Because Claude's knowledge comes from training data with a cutoff date, so claims about recent changes may be outdated or wrong and should be verified against current official sourcesAnswer

    Claude's built-in knowledge stops at its training cutoff, so anything framed as "recent" is exactly where the model is most likely to be stale or to conflate versions of a rule. Time-sensitive claims therefore need verification against the current official source, because a rule may have changed again after the model's knowledge was fixed.

    Source: support.claude.com — Claude's knowledge has a training cutoff date; CCAO-F Exam Guide, Domain 2 (verifying time-sensitive claims)Report a problem with this question

  16. 16. You give Claude a vendor contract and ask questions about its terms. Which technique best grounds Claude's answers in the actual document?

    • A.Ask Claude to answer as briefly as possible so there is less room for error
    • B.Ask the same question three times and accept the most common answer
    • C.Ask the questions without attaching the contract to test Claude's general knowledge
    • D.Ask Claude to first extract the relevant passages word-for-word, then answer based on those quotes, and confirm the quotes actually appear in the contractAnswer

    Extracting verbatim quotes before answering forces the model to anchor its reasoning in the supplied text, and the quotes give the reviewer a direct verification path — searching the document for each quote instantly reveals fabrication. Anthropic's guidance recommends this quote-first grounding for document tasks because it makes hallucinated terms detectable rather than hidden inside a paraphrase.

    Source: Anthropic, "Reduce hallucinations" — use direct quotes for factual grounding in document tasks (docs.claude.com/en/docs/build-with-claude)Report a problem with this question

  17. 17. Which of the following is the clearest example of a hallucination in a Claude output?

    • A.A summary that is shorter than the original document
    • B.An answer that declines to speculate about missing information
    • C.A plausible-sounding statistic attributed to a study that does not existAnswer
    • D.A response written in a different tone than requested

    A hallucination is fabricated content presented as fact — the model invents specifics, like a statistic and its supporting study, that have no basis in reality or in the provided materials. Brevity, appropriate refusal to speculate, and tone mismatches are style or instruction-following issues, not fabrications, which is the defining property of a hallucination.

    Source: Anthropic docs — definition of hallucinations as fabricated/factually incorrect content (docs.claude.com/en/docs/build-with-claude); CCAO-F Exam Guide, Domain 2Report a problem with this question

  18. 18. What is the guiding principle for deciding how much verification a Claude output needs before it is shared or used?

    • A.Skip verification whenever the output cites any source
    • B.Verify only outputs longer than one page
    • C.Scale the rigor of verification with the stakes: light checks for low-risk internal drafts, rigorous source-checking and human review for external, financial, legal, or compliance usesAnswer
    • D.Apply the same maximal verification process to every output regardless of its use

    Verification effort is a finite resource, so effective practice matches the depth of review to the cost of an error: a wrong brainstorm idea costs nothing, while a wrong compliance statement or customer claim can cause legal and financial harm. Uniform maximal review is wasteful and unsustainable, cited sources can themselves be fabricated, and length is unrelated to risk.

    Source: CCAO-F Exam Guide, Domain 2 (risk-proportional verification before sharing outputs)Report a problem with this question

  19. 19. Claude summarizes a policy document and the summary asserts a detail you cannot find anywhere in the document. What is the correct interpretation?

    • A.Claude may have improved the document by adding useful context, so no action is needed
    • B.The detail must be in the document somewhere, so you should keep it in the summary
    • C.The document itself must be incomplete and should be rewritten to match the summary
    • D.The detail is likely a hallucination introduced by the model and should be removed or verified independently before the summary is usedAnswer

    A summary's claims must be traceable to the source document; a detail with no anchor in the source is unsupported by definition, and models are known to blend outside patterns or invented specifics into summaries. The safe handling is to strip or independently verify the unsupported detail, because keeping unverifiable content silently converts a summary into partially fabricated text.

    Source: Anthropic, "Reduce hallucinations" — ground outputs in provided documents and verify traceability (docs.claude.com/en/docs/build-with-claude); CCAO-F Exam Guide, Domain 2Report a problem with this question

  20. 20. An executive asks for "the key numbers" from a long Claude data analysis to paste into a chat message. Which output format best fits this need?

    • A.A multi-page formatted document with appendices
    • B.A raw JSON dump of every intermediate calculation
    • C.A brief inline response with the key figures clearly labeledAnswer
    • D.A full-length artifact reproducing the entire analysis

    Format should follow the consumption context: the executive needs a small set of labeled figures that can be copied into a chat, so a concise inline response delivers exactly that with zero extraction work. An artifact, a raw JSON dump, or a multi-page document each force the recipient to dig the numbers out, defeating the purpose of the request.

    Source: CCAO-F Exam Guide, Domain 2 (matching output format — artifact vs inline vs structured — to how the output will be used); support.claude.com on ArtifactsReport a problem with this question

Practice questions based on the official Claude Certified Associate – Foundations (CCAO-F) exam guide and Anthropic's public documentation. This is an independent study tool, not affiliated with or endorsed by Anthropic, and does not grant certification. The real exam is 60 questions, 120 minutes, passing at a scaled 720/1000, delivered via Pearson VUE ($99). Official certification page →