Modern biology generates information in forms that would have been inconceivable only a few decades ago. Genomic sequences, expression profiles, molecular structures, interaction networks, imaging data and experimental observations can now be produced at enormous scale. Increasingly, they can also be processed by computational systems capable of finding patterns across quantities of information that no individual researcher could reasonably inspect.

This creates an understandable temptation to regard biological representation principally as a problem of data organisation. If biological information can be standardised, encoded and made machine-readable, the reasoning goes, increasingly powerful computational methods can take care of what follows.

But representing biological data is not necessarily the same thing as representing biology.

The distinction becomes increasingly important as computational biology moves from retrieving and classifying information towards modelling mechanisms, generating hypotheses and proposing biological designs.

A sequence is a representation, but of what?

Consider a DNA sequence.

As computational information, its representation is remarkably straightforward: an ordered series drawn from a small alphabet. That representation has transformed biology precisely because it is standardised, portable and amenable to computation.

Yet sequence alone does not tell us everything we might wish to know about the biological system in which that sequence participates.

The function attributed to a sequence may depend upon organism, cell type, regulatory environment and biological state. A molecular interaction may occur under one experimental condition and disappear under another. An observation may be strongly established, weakly inferred or predicted computationally. Two apparently contradictory statements may both be reasonable because they arose in different biological contexts.

If those distinctions disappear during representation, the resulting dataset may remain perfectly machine-readable while becoming less biologically informative.

The problem is therefore not simply how much information we can encode, but which distinctions need to survive the transition from biological knowledge to computational representation.

Standards have already taken us a long way

This is not a new problem, and computational biology has developed sophisticated approaches to it.

The Synthetic Biology Open Language, for example, provides an ontology-backed, machine-tractable representation for biological designs. SBOL3 was explicitly developed to support biological design information across multiple scales while improving computational accessibility and interoperability. [1]

Systems biology has developed complementary standards for mathematical models. More broadly, reproducible computational biology increasingly depends upon standard formats, annotation, documentation and the ability to reconstruct how a model was assembled and used. [2]

These standards matter precisely because representation determines what can subsequently be exchanged, validated, reproduced and computed.

But no single representation needs to contain everything.

A representation designed to exchange a synthetic biological construct, a mathematical model of a signalling pathway and a record of the evidence supporting a host–pathogen interaction serve different purposes. The important question is whether the representations can preserve and connect the information required for the scientific task being undertaken.

The problem becomes harder with biological reasoning

For many computational tasks, relatively narrow representations work extremely well.

If the objective is sequence classification, the representation may need little explicit information about biological mechanism. Modern representation-learning methods can themselves derive useful numerical representations from sequence, gene-expression data, interactions and scientific text. A 2024 systematic review found that much biological representation learning remained focused on individual data types and downstream predictive applications. [3]

That is valuable computational biology.

It is not, however, identical to explicit biological reasoning.

Suppose instead that we ask why a pathogen succeeds in one host but fails in another; how a particular immune response changes the course of infection; or which intervention might disrupt a biological mechanism without creating another undesirable effect.

Now relationships matter. Context matters. Mechanism matters. Time may matter. The origin and strength of the evidence may matter.

A computational system capable of answering such questions reliably needs access not merely to biological observations, but to enough structure around those observations to distinguish what they mean.

This is where biological representation begins to become a scientific question in its own right.

Evidence belongs close to the biology

There is another dimension that becomes particularly important as biological knowledge is assembled from increasingly heterogeneous sources: provenance.

A statement about a biological interaction is not simply true because it exists in a database. It may derive from a particular experiment, organism, assay or publication. It may subsequently have been reproduced, qualified or contradicted.

Representations used for scientific reasoning should therefore make it possible, where appropriate, to travel backwards from an assertion towards the evidence and processes that produced it.

SBOL has incorporated provenance mechanisms into the representation of the design–build–test–learn cycle, drawing on the W3C PROV model. [1]

This principle extends beyond synthetic biology.

If computational systems increasingly participate in scientific inference, preserving the relationship between claim, context and evidence becomes more important rather than less.

Representation before reasoning

At Noviota, this is why we place representation at the beginning of the sequence:

Representation → Modelling → Reasoning → Design

It is tempting to begin with reasoning, particularly now that increasingly capable artificial-intelligence systems can operate directly over scientific literature, sequences and large biological datasets.

But the conclusions available to a computational system are inevitably shaped by what has been made available to it and how that information has been represented.

A representation that preserves only sequence invites certain kinds of reasoning. One that represents entities and relationships permits others. Adding biological context, mechanistic relationships, evidence and provenance changes the questions that can potentially be asked again.

There will probably never be one universal representation of biology, nor should that necessarily be the objective.

Different scientific questions require different abstractions.

The more useful ambition may be to create representations that are explicit about what they preserve, interoperable where appropriate, and sufficiently inspectable that researchers can understand what has been lost as well as what has been encoded.

That is partly what we are exploring through Sybil, Noviota’s experimental domain-specific language for computational biology. Sybil is not intended to replace established biological standards. It provides a research environment in which we can investigate what happens when biological relationships, context and eventually evidence are treated as elements that can be expressed explicitly and manipulated computationally.

The experiment is as much about discovering the limits of representation as extending them.

Because making biology readable by a computer is relatively easy.

Deciding what the computer needs to know about the biology is considerably harder.

Related reading

This argument sits behind why we are building Sybil. It also underlies our research note on what a genome model needs to represent in order to produce a working virus, where statistical representation of sequence proves sufficient to generate viable biology but not to explain it. Our colleagues at Molecular Precision examine the same tension between representation and reproducibility in the design of nanomedicines.

References

  1. McLaughlin, J.A. et al. (2020). The Synthetic Biology Open Language (SBOL) Version 3: Simplified Data Exchange for Bioengineering. Frontiers in Bioengineering and Biotechnology, 8, 1009.
  2. Tiwari, K. et al. (2021). Reproducibility in systems biology modelling. Molecular Systems Biology, 17(2), e9982.
  3. Yang, Y., Zuo, X., Das, A., Xu, H. & Zheng, W. (2024). Representation Learning of Biological Concepts: A Systematic Review. Current Bioinformatics, 19(1), 61–72.