DNA as Information

7 minute read

Published:

DNA is often called an information molecule.

The phrase is so familiar that it can sound obvious.

But what exactly is informational about DNA?

A DNA molecule is a physical polymer.

It has:

  • mass,
  • charge,
  • shape,
  • chemical bonds.

At the same time, the sequence of its bases matters in a way that invites informational description.

The challenge is to understand both levels at once.

Four Bases

DNA uses four nucleotide bases:

  • adenine,
  • cytosine,
  • guanine,
  • thymine.

We abbreviate them:

[ A,\ C,\ G,\ T ]

A DNA molecule is not merely a string of letters.

The letters are our representation of a molecular sequence.

Still, sequence order is biologically crucial.

Sequence as Pattern

Two DNA molecules can be chemically similar in general structure while differing in base sequence.

That difference can alter:

  • protein sequence,
  • gene regulation,
  • development,
  • phenotype.

The informational significance lies largely in structured variation.

DNA as a Physical Record

DNA preserves heritable differences over time.

Replication carries sequence patterns from cell to cell and generation to generation.

This makes DNA a physical memory system.

It stores traces of biological history.

Shannon Information

DNA can be analyzed with information theory.

A sequence drawn from four possible bases can, under idealized equal probabilities, carry up to:

[ \log_2 4 = 2 ]

bits per base.

Real genomes have:

  • biases,
  • repeats,
  • correlations.

So actual information measures depend on model and scale.

Shannon Information Is Not Biological Function

A random DNA sequence can have high Shannon entropy.

That does not mean it has high biological significance.

Once again:

information quantity ≠ meaning.

Functional biology requires context.

Genes

A gene is often described as a DNA segment associated with a functional product.

But modern genetics complicates the old one-gene-one-protein picture.

Genes can produce:

  • multiple RNA products,
  • multiple protein isoforms.

Regulation matters.

Boundaries can be context-dependent.

Coding DNA

Protein-coding regions contain sequences that are translated into amino-acid sequences.

Triplets of bases called codons correspond to amino acids or stop signals.

This code-like mapping is one of the strongest reasons informational language is useful.

Noncoding DNA

Much of the genome does not directly encode protein sequence.

Noncoding regions can include:

  • regulatory elements,
  • structural regions,
  • noncoding RNA genes.

Some sequences have no known function.

“Noncoding” does not mean “useless.”

Regulatory Information

DNA contains sites that influence:

  • when genes turn on,
  • where,
  • how strongly.

This is information about control rather than protein sequence.

The genome contains both product-related and regulatory structure.

Context Dependence

The same DNA sequence can behave differently in different cells.

A neuron and a liver cell share nearly the same genome.

Their gene-expression profiles differ dramatically.

DNA information is interpreted through cellular context.

Genome Is Not Program in a Simple Sense

The computer-program metaphor is tempting.

But a genome does not contain every detail of organismal structure as explicit instructions.

Development depends on:

  • chemical gradients,
  • cell interactions,
  • physical forces,
  • environment.

DNA participates in a dynamic system.

Blueprint Metaphor

The blueprint metaphor also has limits.

A blueprint maps explicitly to spatial structure.

DNA works indirectly.

It influences molecules that influence processes that generate structure.

The organism is constructed through development, not read directly from a drawing.

Recipe Metaphor

A recipe may be somewhat better.

A recipe specifies processes rather than exact final geometry.

But even this metaphor is incomplete.

Cells contain inherited machinery before the genome is used.

The “cook” is already present.

DNA Requires a Cell

A naked DNA molecule does not become an organism by itself.

It requires:

  • polymerases,
  • ribosomes,
  • membranes,
  • energy systems,
  • regulatory networks.

Information becomes functional only inside an interpreting system.

Genetic Code

The genetic code is specifically the mapping from codons to amino acids.

For example, particular codons correspond to particular amino acids.

This mapping is implemented through molecular machinery involving transfer RNA and enzymes.

It is not merely a table in a textbook.

Degeneracy

Several codons can specify the same amino acid.

This property is called degeneracy of the genetic code.

It creates redundancy.

Some mutations therefore do not change the encoded amino acid.

Universality and Variation

The genetic code is highly conserved across life.

But it is not absolutely universal.

Some organisms and organelles use variant mappings.

This shows that the mapping is historically stabilized rather than uniquely forced by pure chemistry.

DNA as Symbol?

Are bases symbols?

In one sense, yes:

sequence elements participate in a code-like system.

In another sense, they are molecules undergoing chemistry.

The symbolic description is useful at one level.

The molecular description is necessary at another.

Causal Role

DNA does not merely represent.

It participates causally.

Its sequence affects:

  • binding,
  • transcription,
  • replication.

Information in biology is embodied in physical interaction.

Reference?

Does a gene “refer” to a protein the way a word refers to an object?

Not exactly.

The relation is mediated by biochemical machinery.

The mapping is functional and causal rather than linguistic in the human sense.

Biological semantics should be used carefully.

Teleosemantic Perspective

One philosophical approach says biological meaning derives from evolutionary function.

A sequence has significance because selection has historically preserved its role.

This can explain why biological signals can be considered correct or incorrect in a functional sense.

The theory remains debated.

Mutation as Information Change

A mutation changes sequence.

Its consequences depend on location and context.

Possible outcomes include:

  • no effect,
  • altered protein,
  • altered regulation.

A one-base change can be trivial or profound.

Bit-count alone cannot tell us which.

Recombination

Recombination reshuffles inherited sequences.

Evolution therefore does not only change information through mutation.

It also reorganizes existing variation.

New combinations can produce new phenotypes.

Evolution Stores History

Genomes contain traces of:

  • duplication,
  • divergence,
  • viral insertion,
  • selection.

DNA is not a cleanly engineered archive.

It is a historical document full of remnants.

Evolution writes by modification, not by starting from scratch.

Gene Duplication

A gene can be copied.

One copy preserves the original function.

The other can accumulate changes.

Duplication creates room for evolutionary novelty.

Information can expand through replication and divergence.

Junk DNA Debate

The term “junk DNA” has been controversial.

Some noncoding DNA has important functions.

Some may be largely neutral.

The safe conclusion is neither:

“all noncoding DNA is junk”

nor:

“all DNA is functional.”

Evidence varies by region and definition of function.

Information Beyond Sequence

Biological inheritance is not exhausted by DNA sequence.

Cells also inherit aspects of:

  • chromatin state,
  • cytoplasmic organization,
  • organelles.

Developmental and environmental processes matter.

DNA is central, not solitary.

The Philosophical Lesson

Calling DNA information is justified because sequence differences are:

  • stored,
  • copied,
  • interpreted,
  • functionally consequential.

But DNA is not an abstract message floating free of chemistry.

Its informational role exists through a physical system.

Biological information is embodied.

The Next Question

How is this information preserved?

How is DNA copied with high fidelity?

And how does sequence information flow toward RNA and protein?

The next essay turns to:

DNA replication and the central dogma.