Experiment

9 minute read

Published:

An experiment asks nature a question by changing something.

Observe.

Intervene.

Compare.

That intervention distinguishes experiment from passive observation.

If two variables are merely correlated, we may not know which causes which.

But if we deliberately change one factor while controlling others, causal structure becomes easier to identify.

This is why experiment occupies such a central place in science.

Observation vs Experiment

Suppose people who exercise more tend to have better cardiovascular health.

That is an observation.

Maybe exercise helps.

Maybe healthier people exercise more.

Maybe income, diet, age, or genetics influence both.

An experiment attempts to break these ambiguities by assigning or manipulating exposure under controlled conditions.

Intervention helps distinguish association from causation.

The Logic of Control

A controlled experiment compares conditions that differ in the factor of interest.

Ideally:

everything relevant is held constant except one variable.

If outcomes differ systematically, the manipulated factor becomes a plausible cause.

In practice, perfect control is impossible.

The art of experimental design is reducing alternative explanations enough that inference becomes credible.

Independent and Dependent Variables

In simple experiments:

  • the independent variable is manipulated,
  • the dependent variable is measured.

For example:

change temperature, measure reaction rate.

But real experiments can involve many interacting variables.

The terminology is useful, not universal.

Complex experimental design often involves multiple factors, covariates, repeated measures, and hierarchical structure.

Control Groups

A control group provides a comparison baseline.

Suppose a treatment group receives a drug.

A control group may receive:

  • placebo,
  • standard treatment,
  • no intervention.

Without a control, natural recovery or background change could be mistaken for treatment effect.

The control tells us what likely would have happened otherwise.

Counterfactual Reasoning

Causal inference depends on a hidden comparison.

What would have happened to the same unit if the treatment had not been applied?

We cannot usually observe both possibilities for the same person or system simultaneously.

This is the fundamental problem of causal inference.

Experiments approximate the missing counterfactual by comparing equivalent groups.

Randomization

Random assignment is one of the most powerful tools in experimental design.

Participants are assigned to conditions by chance.

This helps distribute both known and unknown confounding variables across groups.

With sufficient sample size and proper implementation, differences in outcome can then be attributed more credibly to the intervention.

Randomization is not magic.

But it greatly strengthens causal inference.

Random Sampling vs Random Assignment

These are different.

Random sampling helps generalize from a sample to a population.

Random assignment helps identify causal effects within the sample.

A study can have one without the other.

This distinction is often overlooked.

Internal validity and population generalization are separate questions.

Blinding

Expectations can influence results.

Participants may report improvement because they know they received treatment.

Researchers may unconsciously interpret ambiguous outcomes differently.

Blinding reduces these effects.

Single-blind, double-blind, and related designs vary by who is unaware of treatment assignment.

Blinding is especially important when outcomes involve subjective judgment.

Placebo Effects

Placebos complicate medical experiments.

Belief, expectation, context, and care can affect symptoms and behavior.

A placebo control helps separate the specific effect of an intervention from contextual effects.

The placebo response is not necessarily “imaginary.”

Psychological processes can cause real physiological changes.

The experimental challenge is identifying which mechanism produced which effect.

Confounding

A confounder influences both the presumed cause and the outcome.

Suppose coffee drinking correlates with a health outcome.

If smokers also tend to drink more coffee, smoking may partly explain the association.

Experiments reduce confounding by controlling exposure.

Observational studies must use design and statistical methods to address it.

Manipulation Checks

Researchers need to verify that the intervention actually changed what it was intended to change.

If an experiment manipulates stress, did stress increase?

If a chemical reaction condition is altered, was the concentration correct?

A manipulation check confirms that the experimental treatment existed as designed.

Without it, a null result may simply mean the manipulation failed.

Replication

One experiment rarely settles a question.

Random fluctuations happen.

Hidden biases occur.

Equipment fails.

Results should be reproduced across:

  • laboratories,
  • populations,
  • conditions,
  • measurement methods.

Replication converts isolated findings into reliable knowledge.

Sample Size

Small experiments can produce unstable estimates.

Effects may look large by chance.

Rare confounders may be unbalanced.

Large samples improve statistical precision, though they do not repair systematic design flaws.

A large bad experiment can be confidently wrong.

Quantity does not replace quality.

Power

Statistical power is the probability that a study will detect an effect of a specified size if that effect is real.

Low-powered studies produce two problems.

They miss real effects.

And when they do find significance, estimated effect sizes can be exaggerated.

Power analysis should ideally influence experimental planning before data collection.

Pre-registration

Researchers can unintentionally bias analysis when they explore many possibilities and report only the successful ones.

Pre-registration records hypotheses, outcomes, and analysis plans before seeing results.

It does not forbid exploration.

It distinguishes confirmatory analysis from post hoc discovery.

Transparency helps prevent storytelling after the fact.

P-Hacking

If researchers test enough analyses, some may cross a significance threshold by chance.

Selecting only those results is called p-hacking.

Examples include:

  • trying many outcomes,
  • stopping data collection strategically,
  • excluding inconvenient cases,
  • testing multiple subgroups,
  • switching models repeatedly.

Good experimental practice controls these researcher degrees of freedom.

Multiple Comparisons

If one test has a 5% false-positive rate, running many independent tests increases the chance of at least one false positive.

Scientists use statistical corrections and confirmatory replication to address this.

Large datasets create more opportunities for accidental patterns.

More data do not automatically mean less false discovery.

Natural Experiments

Not every causal question allows controlled assignment.

We cannot randomly assign people to earthquakes or planets to different histories.

Nature or society may create variation resembling experimental conditions.

These are natural experiments.

Careful causal inference can exploit:

  • policy changes,
  • geographic boundaries,
  • historical accidents,
  • natural disasters,
  • genetic variants.

The intervention is not controlled by the researcher, but the comparison can still be informative.

Field Experiments

Laboratory control can reduce realism.

Field experiments intervene in real-world environments.

Examples include:

  • education programs,
  • ecological manipulations,
  • market interventions,
  • public-health messaging.

Field settings improve ecological validity but often reduce control.

Experimental design is a tradeoff.

Laboratory Experiments

Laboratories allow precise control.

Temperature.

Lighting.

Concentration.

Timing.

Stimulus.

Instrumentation.

This makes mechanisms easier to isolate.

But artificial conditions may limit generalization.

A behavior observed in a laboratory may differ in natural settings.

No environment is epistemically perfect.

Physics Experiments

Physics often produces highly controlled systems.

Particle beams.

Vacuum chambers.

Laser pulses.

Cryogenic environments.

Researchers isolate variables to extraordinary precision.

But even fundamental physics depends on complex apparatus and modeling.

Detector calibration, background estimation, and theory-driven reconstruction remain essential.

Biology Experiments

Biological systems are messier.

Organisms vary.

Environments interact with genes.

Multiple pathways produce the same phenotype.

Experiments therefore require larger samples, controls, randomization, and statistical reasoning.

Complexity does not make biology less scientific.

It changes experimental design.

Experiments Can Fail Quietly

A study can produce clean data and still test the wrong thing.

Perhaps the operational definition was poor.

Perhaps the manipulation affected several mechanisms.

Perhaps the outcome measure was insensitive.

Perhaps the population was unrepresentative.

Experimental rigor begins before statistics.

The question itself must be well designed.

Internal Validity

Internal validity asks whether the experiment correctly identifies the causal effect within the studied setting.

Threats include:

  • confounding,
  • attrition,
  • contamination,
  • measurement error,
  • bias.

A highly internally valid experiment may still have poor external validity.

External Validity

External validity asks whether the result generalizes.

Does a treatment work:

  • in other hospitals?
  • in other countries?
  • across age groups?
  • outside laboratory conditions?
  • over longer periods?

A result can be causally correct in one setting yet fail elsewhere.

Generalization requires evidence too.

Mechanism

Experiments can reveal more than whether X affects Y.

They can reveal how.

Disable a pathway.

Change one component.

Measure intermediate states.

A mechanistic experiment decomposes causal chains.

This often produces deeper explanation than simply detecting an average effect.

Ethics

Some of the most informative experiments would be unethical.

We cannot deliberately expose people to severe harm merely to test causation.

Ethical constraints are not obstacles to science in a negative sense.

They define legitimate research.

Scientists use observational studies, natural experiments, animal models, simulations, and other methods when direct intervention is unacceptable.

Experiment Is Not Supreme in Every Field

Astronomy cannot manipulate galaxies.

Evolutionary history cannot be rerun.

Geology cannot restart Earth.

These fields remain scientific because strong inference can arise from observation, prediction, model comparison, and natural variation.

Experiment is powerful.

It is not the only path to knowledge.

The Ideal Experiment Is Rare

Real experiments contain imperfections.

A good scientific conclusion therefore rarely rests on one ideal study.

Confidence grows when:

  • experiments,
  • observational studies,
  • mechanisms,
  • theory,
  • replication

all point in the same direction.

Evidence is a network.

Intervention and Causality

The deepest power of experiment is intervention.

If changing X changes Y under controlled conditions, causal inference becomes stronger.

The experiment does not merely describe patterns.

It probes the structure generating them.

This is why experimentation transformed modern science.

What Comes After the Experiment?

An experiment is designed around expectations.

If theory is correct, what should happen?

If another theory is correct, what should happen instead?

These expectations are predictions.

Prediction gives scientific ideas the ability to risk failure before the data arrive.

That makes it one of science’s sharpest tests.

Why is prediction so important?