Evidence

10 minute read

Published:

Science runs on evidence.

But evidence is not simply “data.”

A number in a spreadsheet is data.

A fossil is an object.

A detector trace is a signal.

A testimony is a statement.

These become evidence only in relation to a claim.

The same observation can support one hypothesis, weaken another, and leave a third unchanged.

Evidence is therefore not a substance.

It is a relationship between what is observed and what is being inferred.

Data vs Evidence

Suppose a thermometer reads 38.5°C.

That is a measurement.

Does it provide evidence of infection?

Possibly.

But body temperature can rise for multiple reasons.

The measurement becomes evidence only when interpreted alongside a hypothesis and background knowledge.

Likewise in science:

data do not speak for themselves.

They speak through models.

Evidence Is Comparative

Evidence is often strongest when it distinguishes between alternatives.

Suppose two theories both predict an eclipse.

Observing the eclipse supports both.

The observation may therefore provide little help in deciding between them.

But if one theory predicts a measurable timing difference and the other does not, precise timing becomes discriminating evidence.

Good tests create divergence among predictions.

Confirmation

Evidence confirms a hypothesis when the observation is more expected if the hypothesis is true than if it is false or replaced by a competitor.

This does not mean the hypothesis becomes certain.

Confirmation is usually gradual.

One result increases confidence.

Repeated independent results increase it further.

Science often works through accumulation rather than one decisive proof.

Disconfirmation

Evidence can also lower confidence.

If a theory strongly predicts X and careful observation repeatedly finds not-X, the theory faces trouble.

But disconfirmation is not always simple.

Perhaps:

  • the instrument failed,
  • an auxiliary assumption was wrong,
  • the system was outside the theory’s domain,
  • the analysis was flawed.

Real theories are tested together with background assumptions.

The Duhem-Quine Problem

A scientific prediction rarely comes from one hypothesis alone.

It also depends on:

  • instrument assumptions,
  • mathematical approximations,
  • initial conditions,
  • auxiliary theories.

When prediction and observation disagree, logic alone may not tell us which component failed.

This is associated with the Duhem-Quine problem.

It does not make testing impossible.

It explains why scientific revision is often more complex than “one failed prediction kills one theory.”

Strong Evidence

Evidence becomes stronger when it has several properties.

It is:

  • independently replicated,
  • produced by reliable methods,
  • difficult to explain under alternatives,
  • predicted in advance,
  • robust across datasets,
  • supported by multiple methods.

A single dramatic result can be interesting.

Converging evidence is usually more convincing.

Independent Evidence

Suppose five laboratories repeat the same experiment using the same flawed instrument design.

The results may not be truly independent.

Independence matters because shared errors can create false agreement.

The strongest cases often combine genuinely different methods.

For example, dark matter evidence comes from galaxy dynamics, lensing, clusters, CMB structure, and large-scale structure.

Different pathways converge on the same conclusion.

Novel Prediction

Evidence can be especially persuasive when a theory predicts something not used to construct it.

General relativity predicted light deflection and other effects.

Dirac’s theory led to antimatter.

The hot Big Bang framework predicted relic radiation.

A successful novel prediction reduces suspicion that a theory was merely adjusted to fit known data.

Accommodation

A theory may also explain data already known.

This is called accommodation.

Accommodation is not worthless.

A theory that unifies previously disconnected observations can be highly explanatory.

But if a theory can be endlessly modified after every result, its apparent success becomes less impressive.

The balance between fit and risk matters.

Ad Hoc Rescue

Suppose a theory fails.

We can sometimes add a new assumption that saves it.

That may be legitimate.

New entities have sometimes been proposed correctly.

Neptune was inferred from orbital anomalies.

Neutrinos were proposed to explain missing energy in beta decay.

The question is whether the added assumption creates independent predictions.

An ad hoc patch that only protects the theory from one failure is less convincing.

Evidence and Probability

Evidence often changes probability rather than proving certainty.

Before an experiment, several hypotheses may have different prior plausibilities.

After the result, those plausibilities should change.

This is the basic intuition behind Bayesian reasoning.

Evidence matters because it updates belief.

The size of the update depends on how expected the evidence was under each hypothesis.

Bayesian Evidence

Bayes’ theorem formalizes belief updating.

In simplified form:

posterior belief ∝ likelihood × prior belief

The likelihood asks:

How probable is this evidence if the hypothesis is true?

A surprising observation under one theory but expected under another can shift belief strongly.

Bayesian reasoning will receive fuller treatment later.

For now, the central lesson is that evidence is comparative and probabilistic.

Likelihood Is Not Posterior Probability

A common mistake is to confuse:

P(datahypothesis)

with

P(hypothesisdata).

These are not the same.

A test can have a high probability of producing a positive result if a condition is present, while the probability that the condition is present after a positive result still depends on prevalence and alternatives.

Scientific evidence requires careful conditional reasoning.

Statistical Evidence

Statistical tests help quantify whether data are surprising under a null model.

But a small p-value does not directly give:

  • probability the hypothesis is true,
  • probability the result will replicate,
  • size of the effect,
  • scientific importance.

Statistics is part of evidential reasoning, not a replacement for it.

Effect Size

A result can be statistically significant but practically tiny.

Suppose a treatment changes an outcome by 0.1%.

With enough data, the effect may be statistically detectable.

But is it biologically or clinically important?

Evidence should include magnitude, not only detectability.

Replication

Replication is one of the strongest evidential filters.

A surprising result can happen by chance.

A biased analysis can create false confidence.

A real effect should often reappear in new data.

Replication is especially important in fields with many variables, flexible analyses, and noisy measurements.

Reproducibility

A result should also be computationally reproducible where possible.

If other researchers cannot obtain the reported result from the same data and code, something is wrong.

Transparent methods let the community inspect the evidential chain.

Secrecy weakens confidence.

Negative Evidence

Failure to observe something can count as evidence.

But only when the theory predicts that the thing should have been observable.

If a detector is too insensitive, a null result means little.

If a highly sensitive experiment repeatedly sees nothing where a model predicts a strong signal, confidence in the model decreases.

Absence of evidence becomes evidence of absence under the right detection conditions.

Testimony as Evidence

Science relies on testimony.

Most researchers do not personally repeat every experiment they cite.

They trust:

  • published papers,
  • databases,
  • calibration laboratories,
  • instrument teams,
  • statistical analyses.

This means scientific knowledge is partly social.

Testimony can be legitimate evidence when supported by reliable institutions and transparent methods.

Expertise as Evidence

For a non-specialist, expert consensus can itself be evidence.

Not because experts are infallible.

Because expertise often tracks:

  • deeper knowledge,
  • familiarity with evidence,
  • awareness of alternatives,
  • methodological competence.

The rational weight of authority depends on domain relevance, independence, conflicts of interest, and quality of consensus.

Anecdotes

Anecdotes are weak evidence for population-level claims.

One patient recovers after taking a treatment.

One person predicts an event correctly.

One unusual coincidence occurs.

These examples may motivate investigation.

But they cannot easily distinguish causation from chance, regression to the mean, selection bias, or memory distortion.

Stories are psychologically powerful and statistically weak.

Extraordinary Claims

The phrase “extraordinary claims require extraordinary evidence” is often associated with Carl Sagan, though the underlying idea is older.

The principle reflects Bayesian reasoning.

A claim with very low prior plausibility requires stronger likelihood evidence to become credible.

This does not mean unusual claims are forbidden.

It means evidence must overcome the weight of alternatives.

Evidence Can Be Misleading

Evidence is not guaranteed to point correctly.

A dataset may be biased.

An instrument may fail.

A statistical fluctuation may look meaningful.

A confounder may mimic causation.

A fraudulent result may appear convincing.

The possibility of misleading evidence is why scientific confidence should depend on networks of support rather than isolated observations.

Consilience

Consilience occurs when evidence from different domains supports the same conclusion.

Evolution is supported by:

  • fossils,
  • comparative anatomy,
  • genetics,
  • biogeography,
  • observed selection.

Plate tectonics is supported by:

  • seafloor spreading,
  • earthquake patterns,
  • magnetic stripes,
  • GPS measurements.

Consilience is powerful because one hidden error is unlikely to explain all independent lines at once.

Evidence Does Not Interpret Itself

Every evidential claim contains background assumptions.

A spectral line is evidence for hydrogen because we trust:

  • atomic theory,
  • instrument calibration,
  • wavelength measurement,
  • source modeling.

This does not make evidence subjective.

It means scientific knowledge is structured.

Claims support other claims.

Some assumptions are themselves supported by enormous independent evidence.

Degrees of Confidence

Science rarely needs only two categories:

true false.

A better scale includes:

  • speculative,
  • plausible,
  • supported,
  • strongly supported,
  • extremely well established.

Confidence should track evidence.

This is more realistic than pretending every scientific statement has equal certainty.

Evidence and Explanation

A theory may fit data without explaining them deeply.

A flexible curve can fit many points.

A mechanistic model may explain why the pattern occurs.

Evidence can therefore support both predictive accuracy and explanatory structure.

The strongest theories often combine both.

From Evidence to Scientific Structures

Evidence does not float alone.

It is organized through:

  • hypotheses,
  • models,
  • theories.

These terms are often used casually as synonyms.

They are not.

A hypothesis is not simply a “small theory.”

A model is not merely a diagram.

A theory is not an unproven guess.

To understand scientific reasoning, we need to separate their roles.

What is the difference between a hypothesis, a theory, and a model?