Capstone: Evaluate a Scientific Claim

LESSON

Scientific Reasoning and Philosophy of Science

008 30 min intermediate CAPSTONE

Capstone: Evaluate a Scientific Claim

By the end of this lesson, you will be able to...

  • Turn a science-flavored headline into a bounded claim with a risky prediction.

  • Audit its method, measurement chain, causal status, replication evidence, model assumptions, and generalization boundary.

  • Write a calibrated decision memo that says what the evidence supports, what remains uncertain, and what test should come next.

Idea in one sentence: A scientific claim deserves trust in proportion to the pressures it survives, not the confidence of its headline.

Core Insight

Imagine a product review that begins: “Retrieval gating makes AI support answers more reliable.” The team has a chart, a statistically summarized experiment, and a model score. A manager wants a rollout recommendation by Friday. You have one page to answer a practical question: What, exactly, has been established, and what would be reckless to infer?

This capstone turns the track into a reusable claim-review memo. It is not a checklist to perform mechanically. Each section should change the strength, scope, or next action of your conclusion. If a measurement ambiguity weakens the causal claim, say so. If a replication reveals a boundary, preserve the boundary instead of hiding it in a footnote. If a model predicts well but has no intervention evidence, do not grant it causal authority by vocabulary.

The central trade-off is that evaluation improves judgment, but it cannot remove all uncertainty. A useful memo narrows a decision without pretending that a finite study has answered every future question.

The Claim Under Review

Start with the smallest sentence that could guide action. Replace “the system is better” with a structured claim:

For eligible English consumer-support requests, retrieval gating using retriever v3, source snapshot S, and rubric v2 reduces the probability of at least one unsupported factual claim compared with the old retrieval path, while increasing unanswered requests by no more than three percentage points.

This sentence names:

If the team cannot fill these fields, it has a topic or aspiration, not yet a scientific claim.

Stage 1: Make the Hypothesis Risky

Write the expected observation and a defeater. A risky version might be:

In a pre-registered five-day trial with randomized session assignment, the gated path will reduce the unsupported-claim rate by at least four percentage points, while the unanswered rate rises by no more than three points. The claim is weakened if the effect is absent under the fixed rubric, if it disappears on a fresh sample, or if the reduction comes only from refusing difficult requests.

The thresholds are not sacred. They make the decision rule visible before the result. A vague prediction such as “quality will improve” can survive any outcome because it has no clear failure condition.

Check the hypothesis for three traps:

  1. Elastic outcome: Would the team change “quality” after seeing the data?
  2. Hidden denominator: Are refusals, unclear labels, or excluded sessions silently removed?
  3. Post-hoc rescue: Is a subgroup or threshold being selected because it makes the result look better?

An honest memo records the original claim and any later refinement separately. Refinement is learning; rewriting history is not.

Stage 2: Audit the Method

Describe how the comparison could identify the claim. Use a compact design table:

Element What the memo must state Support-assistant example
Unit What receives a condition? A user session, assigned before the answer.
Treatment What changes? Retrieval gate v3 with threshold 0.72.
Control What is the counterfactual path? Old retrieval, same model and prompt.
Assignment How are units allocated? Stable random assignment by session ID.
Time When and for how long? Five weekdays, same source snapshot.
Interference Can one unit affect another? Shared cache, agent learning, or source updates.
Compliance Did assigned units receive the intended path? Log gate decisions, fallback, and failures.

Then draw the competing causal stories. The desired path is:

retrieval gate -> better evidence selection -> fewer unsupported claims

Alternatives include:

question difficulty -> routing choice -> outcome
knowledge-base update -> retrieval quality and outcome
evaluator rubric -> recorded outcome
gate -> more refusals -> fewer evaluated answers

Ask whether the design blocks, measures, or merely assumes away each alternative. Randomization helps with measured and unmeasured pre-treatment differences, but it does not fix post-treatment filtering, interference, or a changing evaluator. A causal conclusion must describe the remaining assumptions.

Stage 3: Trace the Measurement Chain

The outcome is not a transparent fact waiting in the logs. Write the chain:

factual support target -> claim-level definition -> rubric v2
                       -> blinded annotators and source snapshot
                       -> labels, unclear cases, exclusions -> rate

For each link, answer:

Report reliability and validity separately. Two annotators can agree on a rubric that measures citation presence instead of factual support. A new rubric can be more valid while breaking historical comparability. Re-label a bridge sample when the instrument changes, and include refusals as an outcome rather than silently treating them as missing.

The memo should also state resolution limits. “12.0%” is not automatically more accurate than “12%.” Precision in display cannot repair a noisy proxy or incomplete knowledge base.

Stage 4: State the Causal Status

Do not let a difference in means do the work of a causal argument. Use a three-level vocabulary:

To justify the strongest level, discuss the counterfactual, assignment, temporal order, mechanism, confounding, interference, and measurement. The mechanism predicts that better source entailment should improve claims, but mechanism alone does not establish the intervention. A model that predicts labels is not automatically a model that tells the team what to change.

State the estimand in plain language. For example:

The average difference in the probability of at least one unsupported claim if each eligible session in the trial population received the gate versus the old path, with refusals and unclear labels handled according to the predeclared protocol.

This sentence prevents accidental movement between answer-level rates, claim-level rates, treated users, and all future users.

Stage 5: Test Replication and Robustness

Use an evidence ladder rather than a single word such as “validated.”

  1. Reproducibility: Can another analyst recompute the chart from the archived data, code, configuration, rubric, and environment?
  2. Exact replication: Does the same treatment and protocol work on a fresh sample from the same population?
  3. Robustness: Does the conclusion survive predeclared changes in threshold, outcome resolution, missingness handling, or time window?
  4. Conceptual replication: Does the mechanism appear with a different retriever, team, product area, or source set?
  5. External validity: Does the claim travel to languages, products, users, or time periods outside the trial?

Interpret failures diagnostically. A failed repeat may reflect sampling variation, implementation drift, protocol mismatch, measurement change, context dependence, or an original error. One successful replication may still share the same hidden rubric or dataset. The memo should identify which rung has been reached and which rung is being requested by the decision.

If the gate helps only when source coverage is high, that is a generalization boundary and possibly a moderator. If the effect disappears when refusals count as failures, the decision trade-off is central, not secondary. If only one subgroup improves, report heterogeneity rather than an average that implies universal benefit.

Stage 6: Place the Claim Under Theory Pressure

Connect the result to a theory without treating theory as decoration. The working theory is:

Evidence filtering improves factual support because the answer generator is less likely to rely on irrelevant or non-entailing material.

List its auxiliary assumptions: source coverage is adequate, relevance tracks entailment, the rubric measures support, the threshold is stable, and the refusal budget reflects the product's purpose.

Now classify observations:

Ask what the model claims. A support score may predict evaluator labels and guide triage without proving that “support” is a literal internal substance or that changing the score causes better answers. Keep predictive, explanatory, ontological, and intervention claims in separate boxes.

The Decision Memo

Bring the audit together in a compact structure:

Claim

One sentence with population, treatment, comparison, outcome, threshold, and time.

Evidence

Sample size, assignment, measurement version, estimate, uncertainty, missingness, and replication rung.

Causal interpretation

What the design identifies, the main alternative explanations, and the assumptions that remain.

Scope and theory

Where the mechanism appears to work, where it fails, which model claims are supported, and which auxiliary assumptions are under pressure.

Decision

Choose one action: deploy, stage a limited rollout, collect a targeted measurement, redesign the experiment, or do not adopt. Tie the choice to stakes, reversibility, and failure cost.

Next test

Name the observation that would most reduce the remaining uncertainty. A good next test is not “collect more data” in the abstract; it might be a bridge-label study, a low-coverage subgroup experiment, an independent conceptual replication, or an intervention that counts refusals as failures.

Worked Conclusion

An appropriately bounded conclusion might read:

In the five-day randomized trial of English consumer-support sessions, retrieval gating v3 produced a lower claim-level unsupported rate under rubric v2 and a three-point increase in unanswered requests. The result is reproducible from the archived analysis and supported by one exact fresh-sample replication. It does not yet establish benefit for low-coverage products, multilingual traffic, or a different evaluator. Because the effect weakens when refusals count as failures, recommend a staged rollout only in high-coverage products, with a preregistered refusal budget and an independent bridge-label check. The next test should compare the gate with a coverage-aware fallback rather than adding another threshold patch.

Notice what the memo does not say. It does not call the gate universally reliable, claim that a score understands truth, or treat the first positive chart as proof of a permanent theory. It still supports a decision: a bounded, reversible rollout with explicit failure signals.

Common Failure Modes

Headline substitution: Repeating “the system is safer” instead of specifying the measured outcome and population.

Method worship: Treating randomization as a guarantee while ignoring changing instruments, interference, or post-treatment exclusions.

Metric reification: Treating a proxy as the phenomenon itself and forgetting the operational definition.

Causal inflation: Moving from association to intervention without naming the counterfactual and alternative explanations.

Replication theater: Calling a different dataset an exact replication, or calling one successful replay universal validation.

Patch accumulation: Adding exceptions whenever a result fails instead of testing a competing mechanism or narrowing scope.

Model overreach: Using predictive performance to justify literal ontology or causal control.

Decision evasion: Listing uncertainty without choosing a proportionate, reversible next action.

Active Checks

Check 1: A persuasive benchmark

A benchmark reports a 10% improvement, but the test set, rubric, and model version are unpublished. The benchmark is run once by its creators. What is the strongest responsible conclusion?

Answer: The result is a preliminary association or performance report under an incompletely auditable measurement chain. Request artifacts, exact replication, and an independent evaluation before making a causal or general deployment claim.

Check 2: A safe decision under uncertainty

The gate's benefit is credible for high-coverage products, uncertain elsewhere, and reversible through staged routing. What decision follows from the memo?

Answer: Stage a limited rollout in the supported boundary with preregistered outcomes, monitor refusals and claim support, and run a targeted test in low-coverage products. Uncertainty narrows the action; it does not force either blind rollout or total inaction.

Transfer Practice

Choose one claim from a paper, benchmark, health article, dashboard, or technical proposal. Write a one-page memo using these headings:

  1. Claim and risky prediction.
  2. Unit, treatment, control, and assignment.
  3. Measurement chain and proxy risk.
  4. Descriptive, predictive, and causal status.
  5. Replication, robustness, and generalization.
  6. Theory, model, and auxiliary assumptions.
  7. Decision, boundary, and next test.

End with this sentence frame:

“The evidence supports ___ for ___ under . It does not yet justify . The next observation that would most change my decision is ___.”

That sentence is the durable artifact of the track: neither science as a trust costume nor skepticism as a reflex, but a claim whose strength is matched to the evidence it has survived.

Resources

Key Takeaways

PREVIOUS Models, Realism, and Useful Fiction