Capstone: Evaluate a Scientific Claim
LESSON
Capstone: Evaluate a Scientific Claim
By the end of this lesson, you will be able to...
Turn a science-flavored headline into a bounded claim with a risky prediction.
Audit its method, measurement chain, causal status, replication evidence, model assumptions, and generalization boundary.
Write a calibrated decision memo that says what the evidence supports, what remains uncertain, and what test should come next.
Idea in one sentence: A scientific claim deserves trust in proportion to the pressures it survives, not the confidence of its headline.
Core Insight
Imagine a product review that begins: “Retrieval gating makes AI support answers more reliable.” The team has a chart, a statistically summarized experiment, and a model score. A manager wants a rollout recommendation by Friday. You have one page to answer a practical question: What, exactly, has been established, and what would be reckless to infer?
This capstone turns the track into a reusable claim-review memo. It is not a checklist to perform mechanically. Each section should change the strength, scope, or next action of your conclusion. If a measurement ambiguity weakens the causal claim, say so. If a replication reveals a boundary, preserve the boundary instead of hiding it in a footnote. If a model predicts well but has no intervention evidence, do not grant it causal authority by vocabulary.
The central trade-off is that evaluation improves judgment, but it cannot remove all uncertainty. A useful memo narrows a decision without pretending that a finite study has answered every future question.
The Claim Under Review
Start with the smallest sentence that could guide action. Replace “the system is better” with a structured claim:
For eligible English consumer-support requests, retrieval gating using retriever v3, source snapshot S, and rubric v2 reduces the probability of at least one unsupported factual claim compared with the old retrieval path, while increasing unanswered requests by no more than three percentage points.
This sentence names:
- the population: eligible English consumer-support requests;
- the intervention: a specific gate and retriever version;
- the comparison: the old retrieval path;
- the outcome: at least one unsupported factual claim;
- the constraint: a refusal or unanswered-request budget;
- and the time and instrument context, which must be added in the actual memo.
If the team cannot fill these fields, it has a topic or aspiration, not yet a scientific claim.
Stage 1: Make the Hypothesis Risky
Write the expected observation and a defeater. A risky version might be:
In a pre-registered five-day trial with randomized session assignment, the gated path will reduce the unsupported-claim rate by at least four percentage points, while the unanswered rate rises by no more than three points. The claim is weakened if the effect is absent under the fixed rubric, if it disappears on a fresh sample, or if the reduction comes only from refusing difficult requests.
The thresholds are not sacred. They make the decision rule visible before the result. A vague prediction such as “quality will improve” can survive any outcome because it has no clear failure condition.
Check the hypothesis for three traps:
- Elastic outcome: Would the team change “quality” after seeing the data?
- Hidden denominator: Are refusals, unclear labels, or excluded sessions silently removed?
- Post-hoc rescue: Is a subgroup or threshold being selected because it makes the result look better?
An honest memo records the original claim and any later refinement separately. Refinement is learning; rewriting history is not.
Stage 2: Audit the Method
Describe how the comparison could identify the claim. Use a compact design table:
| Element | What the memo must state | Support-assistant example |
|---|---|---|
| Unit | What receives a condition? | A user session, assigned before the answer. |
| Treatment | What changes? | Retrieval gate v3 with threshold 0.72. |
| Control | What is the counterfactual path? | Old retrieval, same model and prompt. |
| Assignment | How are units allocated? | Stable random assignment by session ID. |
| Time | When and for how long? | Five weekdays, same source snapshot. |
| Interference | Can one unit affect another? | Shared cache, agent learning, or source updates. |
| Compliance | Did assigned units receive the intended path? | Log gate decisions, fallback, and failures. |
Then draw the competing causal stories. The desired path is:
retrieval gate -> better evidence selection -> fewer unsupported claims
Alternatives include:
question difficulty -> routing choice -> outcome
knowledge-base update -> retrieval quality and outcome
evaluator rubric -> recorded outcome
gate -> more refusals -> fewer evaluated answers
Ask whether the design blocks, measures, or merely assumes away each alternative. Randomization helps with measured and unmeasured pre-treatment differences, but it does not fix post-treatment filtering, interference, or a changing evaluator. A causal conclusion must describe the remaining assumptions.
Stage 3: Trace the Measurement Chain
The outcome is not a transparent fact waiting in the logs. Write the chain:
factual support target -> claim-level definition -> rubric v2
-> blinded annotators and source snapshot
-> labels, unclear cases, exclusions -> rate
For each link, answer:
- What construct is the team trying to represent?
- What counts as an observation?
- What is the unit: answer, sentence, or atomic claim?
- What evidence must entail the claim, including time and scope?
- How are ambiguous and missing cases recorded?
- Which instrument version and annotator protocol produced the label?
Report reliability and validity separately. Two annotators can agree on a rubric that measures citation presence instead of factual support. A new rubric can be more valid while breaking historical comparability. Re-label a bridge sample when the instrument changes, and include refusals as an outcome rather than silently treating them as missing.
The memo should also state resolution limits. “12.0%” is not automatically more accurate than “12%.” Precision in display cannot repair a noisy proxy or incomplete knowledge base.
Stage 4: State the Causal Status
Do not let a difference in means do the work of a causal argument. Use a three-level vocabulary:
- Descriptive: The gated group had a measured rate of 12% and the control group 18%.
- Associational or predictive: Under this assignment and population, sessions routed to the gate were associated with fewer unsupported claims, and a similar future population may show a lower rate.
- Causal: For comparable eligible sessions, routing through the gate rather than the old path would lower the probability of an unsupported claim under the specified treatment and measurement versions.
To justify the strongest level, discuss the counterfactual, assignment, temporal order, mechanism, confounding, interference, and measurement. The mechanism predicts that better source entailment should improve claims, but mechanism alone does not establish the intervention. A model that predicts labels is not automatically a model that tells the team what to change.
State the estimand in plain language. For example:
The average difference in the probability of at least one unsupported claim if each eligible session in the trial population received the gate versus the old path, with refusals and unclear labels handled according to the predeclared protocol.
This sentence prevents accidental movement between answer-level rates, claim-level rates, treated users, and all future users.
Stage 5: Test Replication and Robustness
Use an evidence ladder rather than a single word such as “validated.”
- Reproducibility: Can another analyst recompute the chart from the archived data, code, configuration, rubric, and environment?
- Exact replication: Does the same treatment and protocol work on a fresh sample from the same population?
- Robustness: Does the conclusion survive predeclared changes in threshold, outcome resolution, missingness handling, or time window?
- Conceptual replication: Does the mechanism appear with a different retriever, team, product area, or source set?
- External validity: Does the claim travel to languages, products, users, or time periods outside the trial?
Interpret failures diagnostically. A failed repeat may reflect sampling variation, implementation drift, protocol mismatch, measurement change, context dependence, or an original error. One successful replication may still share the same hidden rubric or dataset. The memo should identify which rung has been reached and which rung is being requested by the decision.
If the gate helps only when source coverage is high, that is a generalization boundary and possibly a moderator. If the effect disappears when refusals count as failures, the decision trade-off is central, not secondary. If only one subgroup improves, report heterogeneity rather than an average that implies universal benefit.
Stage 6: Place the Claim Under Theory Pressure
Connect the result to a theory without treating theory as decoration. The working theory is:
Evidence filtering improves factual support because the answer generator is less likely to rely on irrelevant or non-entailing material.
List its auxiliary assumptions: source coverage is adequate, relevance tracks entailment, the rubric measures support, the threshold is stable, and the refusal budget reflects the product's purpose.
Now classify observations:
- A changed rubric is a measurement issue until a bridge sample shows otherwise.
- Failure in low-coverage products may be a scope boundary.
- Persistent failure after stable exact and conceptual replications is pressure on the mechanism.
- A competing approach that improves support and refusals across coverage levels deserves comparison, not another exception added to the old story.
Ask what the model claims. A support score may predict evaluator labels and guide triage without proving that “support” is a literal internal substance or that changing the score causes better answers. Keep predictive, explanatory, ontological, and intervention claims in separate boxes.
The Decision Memo
Bring the audit together in a compact structure:
Claim
One sentence with population, treatment, comparison, outcome, threshold, and time.
Evidence
Sample size, assignment, measurement version, estimate, uncertainty, missingness, and replication rung.
Causal interpretation
What the design identifies, the main alternative explanations, and the assumptions that remain.
Scope and theory
Where the mechanism appears to work, where it fails, which model claims are supported, and which auxiliary assumptions are under pressure.
Decision
Choose one action: deploy, stage a limited rollout, collect a targeted measurement, redesign the experiment, or do not adopt. Tie the choice to stakes, reversibility, and failure cost.
Next test
Name the observation that would most reduce the remaining uncertainty. A good next test is not “collect more data” in the abstract; it might be a bridge-label study, a low-coverage subgroup experiment, an independent conceptual replication, or an intervention that counts refusals as failures.
Worked Conclusion
An appropriately bounded conclusion might read:
In the five-day randomized trial of English consumer-support sessions, retrieval gating v3 produced a lower claim-level unsupported rate under rubric v2 and a three-point increase in unanswered requests. The result is reproducible from the archived analysis and supported by one exact fresh-sample replication. It does not yet establish benefit for low-coverage products, multilingual traffic, or a different evaluator. Because the effect weakens when refusals count as failures, recommend a staged rollout only in high-coverage products, with a preregistered refusal budget and an independent bridge-label check. The next test should compare the gate with a coverage-aware fallback rather than adding another threshold patch.
Notice what the memo does not say. It does not call the gate universally reliable, claim that a score understands truth, or treat the first positive chart as proof of a permanent theory. It still supports a decision: a bounded, reversible rollout with explicit failure signals.
Common Failure Modes
Headline substitution: Repeating “the system is safer” instead of specifying the measured outcome and population.
Method worship: Treating randomization as a guarantee while ignoring changing instruments, interference, or post-treatment exclusions.
Metric reification: Treating a proxy as the phenomenon itself and forgetting the operational definition.
Causal inflation: Moving from association to intervention without naming the counterfactual and alternative explanations.
Replication theater: Calling a different dataset an exact replication, or calling one successful replay universal validation.
Patch accumulation: Adding exceptions whenever a result fails instead of testing a competing mechanism or narrowing scope.
Model overreach: Using predictive performance to justify literal ontology or causal control.
Decision evasion: Listing uncertainty without choosing a proportionate, reversible next action.
Active Checks
Check 1: A persuasive benchmark
A benchmark reports a 10% improvement, but the test set, rubric, and model version are unpublished. The benchmark is run once by its creators. What is the strongest responsible conclusion?
Answer: The result is a preliminary association or performance report under an incompletely auditable measurement chain. Request artifacts, exact replication, and an independent evaluation before making a causal or general deployment claim.
Check 2: A safe decision under uncertainty
The gate's benefit is credible for high-coverage products, uncertain elsewhere, and reversible through staged routing. What decision follows from the memo?
Answer: Stage a limited rollout in the supported boundary with preregistered outcomes, monitor refusals and claim support, and run a targeted test in low-coverage products. Uncertainty narrows the action; it does not force either blind rollout or total inaction.
Transfer Practice
Choose one claim from a paper, benchmark, health article, dashboard, or technical proposal. Write a one-page memo using these headings:
- Claim and risky prediction.
- Unit, treatment, control, and assignment.
- Measurement chain and proxy risk.
- Descriptive, predictive, and causal status.
- Replication, robustness, and generalization.
- Theory, model, and auxiliary assumptions.
- Decision, boundary, and next test.
End with this sentence frame:
“The evidence supports ___ for ___ under . It does not yet justify . The next observation that would most change my decision is ___.”
That sentence is the durable artifact of the track: neither science as a trust costume nor skepticism as a reflex, but a claim whose strength is matched to the evidence it has survived.
Resources
- [ARTICLE] Scientific Method - Stanford Encyclopedia of Philosophy - Focus: hypotheses, testing, evidence, and fallibility.
- [ARTICLE] Causal Inference - Stanford Encyclopedia of Philosophy - Focus: interventions, counterfactuals, and causal interpretation.
- [ARTICLE] Replication - Stanford Encyclopedia of Philosophy - Focus: replication, robustness, and how failures should be interpreted.
- [ARTICLE] Scientific Realism - Stanford Encyclopedia of Philosophy - Focus: what models and theories justify us in believing.
Key Takeaways
- A scientific claim review traces one chain: risky hypothesis, design, measurement, causal status, replication, theory, model, limits, and decision.
- The strongest conclusion is not the most confident sentence; it is the most specific sentence the evidence can support.
- Uncertainty should produce a bounded, proportionate next action, not a vague disclaimer or an unjustified rollout.
- A good memo preserves defeaters and boundaries so the claim can be tested again rather than protected by rhetoric.
← Back to Scientific Reasoning and Philosophy of Science