Blog
The Alphabet, the Readout, the Lock, and the Movie: How AI Is Probing the Genome
The Alphabet, the Readout, the Lock, and the Movie: How AI Is Probing the Dark Matter of the Genome
There is something quietly beautiful about the recent history of genomics: almost every major technique was born because the previous one could not answer the next question.
You could tell this story as a parade of machines, protocols, and acronyms: Sanger, NGS, RNA-seq, ATAC-seq, single-cell RNA-seq, Hi-C, live-cell imaging, sequence-to-function models, AlphaGenome.
But that would miss the real plot.
For the last fifty years, we have been stacking layers of information on top of one another to understand one deceptively simple question:
How does a cell know what to do?
Same DNA. Same alphabet. And yet a neuron, a hepatocyte, a lymphocyte, and a cancer cell clearly do not live in the same world.
Modern biology begins in that gap: the gap between what is written and what is actually performed.
I see that gap almost every day. A large part of my work consists of trying to understand why a cell has tipped into a different state — and sometimes searching through gigabytes of sequence data for the exact line that might explain that shift. This essay starts there: not from an abstract fascination with artificial intelligence, but from a field where the question “what did this cell do with its genome?” has a name, a medical file, and sometimes a family waiting for an answer.
1. The Alphabet: Sequence and Its Illusions
Sequencing gives us the alphabet.
In 1977, Frederick Sanger and colleagues published the chain-termination method for DNA sequencing.1 At the time, reading a few hundred letters of DNA was already a technical achievement. Today, we handle huge genomic files almost casually in a terminal, as if three billion bases were just a slightly overweight PDF.
Then came the Human Genome Project. In June 2000, the international consortium announced a first draft of the human genome. In April 2003, it announced a sequence considered “essentially complete” for the technologies of the time, covering around 90% of the human genome, while still leaving difficult regions unresolved.2
Next-generation sequencing — NGS — really arrived after that, in the mid-2000s, with technologies such as 454 pyrosequencing.3 That changed the scale of the problem: we were no longer reading a single human reference genome. We were moving toward hundreds, then thousands, then millions of genomes.
And in 2022, the Telomere-to-Telomere consortium finally filled in the major remaining gaps with T2T-CHM13, the first complete human genome sequence from end to end, resolving many repetitive and structurally complex regions that had resisted earlier technologies.4
But here is the trap: having the alphabet does not mean you understand the text.
Humans and chimpanzees share the vast majority of their DNA letters. That does not explain why one writes blog posts about genomics and the other probably has the good sense not to waste its time on the internet.
The same problem exists inside our own bodies. With a few important exceptions — somatic mutations, mosaicism, immune-cell rearrangements, cancer — the cells of an individual share the same inherited nuclear genome, the same constitutional or germline sequence. And yet they have radically different fates.
A neuron does not do the job of a hepatocyte. A retinal cell does not behave like a muscle cell.
A cancer cell can even betray the tissue it came from entirely. I see this regularly in molecular tumor work: by the time we characterize some tumors, they barely “look” like their organ of origin anymore.
Sequence tells us what is written.
It does not, by itself, tell us what is being used.
And there is another complication: only a small fraction of the human genome directly codes for proteins. People often cite a figure around 1.5–2%, although the exact number depends on what you count: exons, strictly coding sequences, annotated transcripts, or regions actually translated into proteins.5
The rest was once too casually thrown into the category of “junk DNA.”
That was a vocabulary mistake. And in science, vocabulary mistakes often become thinking mistakes.
Non-coding regions are not an empty desert. They contain enhancers, promoters, insulators, non-coding RNAs, repetitive sequences, transposable elements, three-dimensional architectures, and many regions we still do not understand well.
Not all of it is functional. Not all of it is a master control room. But a major part of gene regulation happens there.
That is why people sometimes call it the dark matter of the genome.
Not because it is magical.
Because we know it matters, but we do not always know how to measure it, interpret it, and connect it cleanly to disease.
2. The Readout: Expression and Its Background Noise
The next layer is the readout.
Which genes are copied into RNA? In which tissue? At what time? At what level? In which cell?
From Northern blots in the 1970s to DNA microarrays in the 1990s, and then RNA-seq from the late 2000s onward, the question became: what is the cell actually reading from its genome?6
RNA-seq changed the field because it turned expression into a massively measurable signal. We were no longer following one favorite gene under a magnifying glass. We could look at the entire transcriptome, with its isoforms, expression levels, splice variants, surprises, and false friends.
Projects such as GTEx mapped gene expression and genetic regulatory effects across many human tissues.7 The Human Cell Atlas goes even further, aiming to map the human body at single-cell scale, with much finer cellular and spatial resolution.8
We learned that the difference between cells is not only in the alphabet. It is in what gets read, amplified, ignored, repressed, or remixed.
But here again, there is a classic trap: expression is not causality.
A gene that is highly expressed in a tumor, for example, may be the driver of the disease — or merely a symptom of it. Expression alone does not settle the question. Usually, we need to cross it with other layers, and often with other patients, before we dare to answer.
3. The Lock: Chromatin Accessibility
To understand why some genes are read while others remain silent, we need to move into a more physical layer: chromatin.
Inside the nucleus, DNA is not a naked thread floating politely in space. It is wrapped around histone proteins, compacted into nucleosomes, folded, looped, constrained, organized.
A DNA region may contain a beautiful enhancer on paper. But if that region is packed too tightly, the cellular machinery cannot reach it.
That is where the idea of the lock comes in.
An open region is a region the cell can potentially use. A closed region is physically less accessible.
To map these open regions, scientists first used methods such as DNase-seq, then ATAC-seq, introduced in 2013, which identifies accessible chromatin regions and sequences them afterward.9
A decent analogy would be this: imagine walking through a house at night with a flashlight. You can only inspect the rooms whose doors are open. But an open door does not mean anything is actually happening inside the room.
This is probably one of the most important traps in modern epigenomics.
Accessibility does not mean activity.
A region can be open, prepared, primed, but inactive. People sometimes speak of “primed” or “poised” enhancers: ready, but not yet switched on — a distinction that certain chromatin marks help us make.10
In other words:
the door is open;
the room is available;
but the party has not started.
This nuance matters enormously for AI in genomics. If we naively feed a model sequence, expression, accessibility, histone marks, and clinical data as if they were independent variables, we get a beautiful statistical soup.
Accessibility partly depends on sequence. Expression partly depends on accessibility. And the whole system depends on cell type, timing, environment, measurement technology, batch effects, and experimental noise.
Stacking layers is not enough.
We have to understand how they relate.
4. The Movie: Life in Motion
The alphabet, the readout, and the lock are often snapshots.
Beautiful snapshots. High-resolution snapshots. Sometimes incredibly precise snapshots.
But snapshots all the same.
Life is a movie.
A cell does not merely have a state. It changes state. It migrates. It divides. It dies. It differentiates. It responds to stress. It sometimes hesitates between two possible fates. It branches.
Modern microscopy — not my specialty, but still — time-lapse imaging, and light-sheet microscopy add that dynamic layer. We no longer only observe a cell at one given moment; we follow its trajectory through time.11
A cell changing fate, a cell population reorganizing, a tumor invading, an embryo taking shape: all of this belongs to motion.
And this movie is becoming readable by AI as well.
Image segmentation, cell tracking, trajectory prediction, rare-event detection: we can imagine models that no longer read only letters or expression matrices, but behaviors.
The cell becomes a character.
And biology starts to look less like a static encyclopedia and more like a 35 mm film reel.
5. Why Genomics Is Almost an Ideal Playground for AI
What makes genomics so interesting for artificial intelligence is not only the massive volume of data. We already knew that.
It is the structure of the data.
Sequence has coordinates. Expression is attached to genes, transcripts, cells, tissues. Accessibility can be aligned to precise genomic positions. Variants can be placed on a reference and their effects predicted or confirmed. Slide images can be linked to cellular phenotypes, expression data, and other molecular layers.
In short: a large part of biology becomes computable.
But not all of it.
The temptation would be to believe that we can simply feed everything into the models — sequence, RNA-seq, ATAC-seq, ChIP-seq, Hi-C, imaging, clinical data, literature — during training and then inference, and let the machine find the truth.
That is the slightly naive version of AI in biology.
The other version, less glamorous, is much more interesting.
It asks:
- what data enters the model, and why?
- what output does the model produce?
- what uncertainty is hidden?
- which tissue is represented?
- which cell type is actually being modeled?
- which population was used for training?
- what experiment could invalidate the prediction?
This is the strong idea developed by Li Lei in “From Genomics to AI Biology”: the challenge is no longer merely to use AI tools, because by 2026 everyone is doing that, but to think in AI systems.12
A biologist using a tool might ask:
“Can you rank this variant for me?”
A biologist thinking in systems asks:
“What data is this ranking based on? In which biological context is it valid? What competing hypotheses remain possible? What experiment should I do next?”
That difference changes everything.
A gene is not “important” in the abstract. It is important in a tissue, a developmental stage, a cell, a species, an environment, a genomic architecture, an evolutionary history.
A variant is not simply “pathogenic” or “benign” like a label slapped onto a box. It has an allele frequency, a possible effect, uncertainty, biological plausibility, family segregation, penetrance, and clinical context.
A regulatory region is not simply “open.” It may be accessible, bound by a transcription factor, conserved, active, silent, redundant, specific to one cell type, or completely irrelevant in the patient’s tissue.
AI can accelerate discovery.
It does not make biology less biological.
6. The Age of Sequence-to-Function Models
For a long time, predicting the impact of a variant mostly meant looking at conservation, position, protein effect, or a few local annotations.
That was useful.
But very insufficient for the non-coding genome.
The new generation of models is trying to do something different: learn the regulatory grammar that links sequence to function.
These are often called sequence-to-function or sequence-to-activity models.
The idea is easy to state and painfully hard to execute:
If I give a model a DNA sequence, can it predict what that sequence does in a cell?
Can it predict expression? Accessibility? Histone marks? Transcription factor binding? Splicing? Three-dimensional contacts?
Models such as Enformer showed that deep learning could improve gene expression prediction from DNA sequence by integrating long-range regulatory interactions.13
Borzoi then pushed this logic further by learning to predict tissue- and cell-type-specific RNA-seq coverage profiles directly from DNA sequence.14
Then AlphaGenome arrived.
Announced by Google DeepMind in 2025 and published in Nature on January 28, 2026, AlphaGenome pushes the paradigm much further. It takes up to one million DNA bases as input and predicts thousands of functional signals, sometimes at single-base resolution.1516
It sees wide and fine at the same time.
That is the old trade-off in genomic modeling. Either you look at a small region with high resolution, or you look at a large region with coarser resolution. AlphaGenome tries to reduce that compromise.
Among other things, it predicts:
- gene expression;
- transcription initiation;
- chromatin accessibility;
- certain histone marks;
- transcription factor binding;
- splice-site usage — and it may eventually absorb part of what specialized tools such as SpliceAI currently do;
- splice junctions;
- chromatin contact maps.
One thing worth noting: AlphaGenome already has an active community of testers on the official forum — scientists, engineers, and curious non-specialists — and some of the discussions there are genuinely useful.
7. Virtual Mutagenesis: Testing a Letter Without Touching the Bench
The main practical appeal of these models is virtual mutagenesis.
Take a sequence.
Change one letter.
Ask the model: what changes?
Does expression of a gene go down? Does a splice site disappear? Does a region become less accessible? Does an enhancer lose predicted activity? Does a 3D interaction seem to shift?
This is obviously attractive for rare diseases, cancer, non-coding variants, and all the cases where sequencing finds something suspicious without clearly telling us what to do with it.
In clinical genetics, the nightmare has a name: VUS — variant of uncertain significance.
You find a variant. It is rare. It is plausible. It is in the right gene. But you do not know whether it truly explains the patient’s phenotype.
And when that variant falls in a non-coding region, the difficulty climbs again. This is not a marginal situation: around 98% of human genetic variation is estimated to lie in regions that do not code for proteins — almost the mirror image of the 1.5–2% coding fraction mentioned earlier.17
This is precisely the kind of problem MobiDeep is working on: an AI meta-score for prioritizing non-coding variants, developed by the MoBiDiC team at Montpellier University Hospital, a group I am part of, alongside David Baux and the rest of the team.18
Specialized tools have already started to change practice. SpliceAI predicts potential splicing effects from primary sequence.19 AlphaMissense focuses on missense variants in coding regions and predicts whether an amino-acid substitution is more likely to be pathogenic or benign.20
AlphaGenome and its successors promise a more integrated and practical approach. Instead of having one tool per molecular layer, we query a system capable of predicting several molecular layers at once.
But this is where we have to stay sober — and I am saying this as much to myself as to anyone else.
A prediction is not proof.
A score is not a diagnosis.
A model can prioritize a variant, suggest a mechanism, guide a biological experiment, speed up a case discussion, but it will never replace validation.
To turn a prediction into robust knowledge, we still need experimental and clinical evidence: family studies, precise phenotyping, functional assays, biological coherence.
For this use case, AI does not close the file.
It tells us where to dig.
And sometimes that is already huge.
8. Scientific AI Agents
Another evolution is arriving very, very fast: scientific AI agents.
The idea is no longer to have only a model that predicts a score, but a system capable of planning and executing a sequence of actions, almost like a bioinformatician.
For example:
- retrieve a sequence;
- annotate variants;
- query public databases;
- read the literature;
- predict an effect;
- model a protein structure;
- propose an experiment;
- generate an interpretable report.
Frameworks such as NVIDIA BioNeMo Agent Toolkit are moving in this direction.21
This is very promising.
But it is also exactly where the hype machine becomes dangerous.
The word “agent” quickly gives the impression that a tiny autonomous researcher lives inside the computer, wearing a virtual lab coat, ready to do science while we drink coffee.
Not really.
An AI agent is mostly an orchestrator. It calls tools, chains steps together, manipulates outputs, and can draft a report. It can also misinterpret a result, use the wrong database version, hallucinate a reference, ignore a confounder, or confidently accelerate a bad hypothesis.
Automating a workflow is not the same as understanding a phenomenon.
An AI agent can speed up an analysis.
It can also speed up an error.
So where are the guardrails? Who checks the sources? Who controls database versions? Who knows whether the model is valid in this tissue? Who detects the batch effect? Who sees that the hypothesis is biologically absurd despite a beautiful score?
This is where the biologist and the physician remain indispensable.
Update — June 30, 2026: Claude Science
Anthropic has just announced Claude Science, an AI workbench designed for researchers. The speed at which this field is moving is honestly wild.22
The idea is to stop forcing biologists — or researchers more broadly — to jump between a dozen tools: PubMed, Jupyter, R, a cluster terminal, and half a dozen databases, each with its own schema. Instead, Claude Science brings them into one working environment. A general-purpose agent coordinates more than 60 preconfigured skills and connectors for genomics, single-cell biology, proteomics, structural biology, and cheminformatics. It can query databases such as UniProt, PDB, Ensembl, Reactome, ClinVar, ChEMBL, or GEO in natural language, instead of forcing users to juggle their respective APIs.22
Three details caught my attention because they speak directly to the concerns I raised above about AI agents.
First, every figure or result is delivered with the exact code, environment, and full history used to generate it. That means the result can be validated and reproduced months later, instead of being taken on trust.22
Second, a “reviewer” agent runs in parallel to check citations and calculations, flag numbers that cannot be traced back to their source, and identify figures that do not match the code used to generate them. That is a partial answer to the question I asked earlier: who checks the sources, and who catches the error the agent just accelerated?22
The system also builds on NVIDIA’s BioNeMo Agent Toolkit, already mentioned above, to connect with models such as Evo 2, Boltz-2, or OpenFold3. You can see the ecosystem taking shape: not as a single isolated tower, but as a network of tools that plug into one another.22
One early example cited by Anthropic: a molecular epidemiology lab at UCSF says it can now explore how thousands of small-effect germline variants combine in glioma susceptibility in about one-tenth of the usual time, with results independently validated by the team itself.22
Does this close the case opened earlier? No.
Traceability and automated review answer part of the problem, not its core. An agent that cites sources correctly can still start from a biologically absurd hypothesis. Speed still does not replace judgment.
But it does confirm one thing: scientific AI agents are already moving into laboratories.
9. Augmented, Not Replaced
Does all of this deserve the word “augmented”?
Yes.
But not for the reason people usually mean.
AI does not augment the biologist because it is “more intelligent.” It augments the biologist because it extends what is already explicit, structured, measurable, and repeatable.
It reads faster. It compares more broadly. It ranks more systematically. It detects patterns no human can track across an entire genome. It can perform in seconds a first pass over hypotheses that a team might need days to triage.
But it remains weak in one crucial area: tacit knowledge.
Tacit knowledge is the beautiful concept associated with Michael Polanyi, often summarized by the sentence: “We can know more than we can tell.”23
It is the lab technician who sees that a slide “doesn’t look right” before being able to explain why.
It is the clinician who understands that a detail in the face, family story, or disease trajectory changes the entire interpretation.
It is the experienced researcher who distrusts a signal because they saw a similar artifact ten months earlier in another experiment.
It is embodied, situated, rough-edged experience.
This knowledge is not always in databases. It is not always written in papers. It is not always formalized as variables.
So it is not easily learned by a model.
AI is excellent at explicit computation.
But real scientific practice also lives in the implicit.
10. What AI Actually Changes
AI does not magically decode the dark matter of the genome.
It does something more useful:
it turns part of that darkness into testable hypotheses.
It says: this mutation may disrupt a splice site. This non-coding region may alter the expression of this gene in this cell type. This sequence looks like an active enhancer in this biological context. This hypothesis may be worth an experiment.
That is certainly less spectacular than saying “AI understands the genome.”
But it is much more powerful.
Because science does not move forward only when we get answers.
It moves forward when we ask better questions.
The next step will not be one lonely model that “solves” the human genome.
It will look more like a loop:
structured data → predictive model → hypothesis → experiment → validation → back into the model.
A dialectical loop where sequence, expression, clinic, literature, and human intuition respond to one another.
This essay is the first in a series. Next, I want to explore foundation models for tabular data, architectures such as Transformers, and maybe even… JEPA.
References and Further Reading
Sanger F., Nicklen S., Coulson A.R. “DNA sequencing with chain-terminating inhibitors.” PNAS, 1977. https://doi.org/10.1073/pnas.74.12.5463 ↩︎
National Human Genome Research Institute. “Human Genome Project Fact Sheet.” https://www.genome.gov/about-genomics/educational-resources/fact-sheets/human-genome-project ↩︎
Margulies M. et al. “Genome sequencing in microfabricated high-density picolitre reactors.” Nature, 2005. https://doi.org/10.1038/nature03959 ↩︎
Nurk S. et al. “The complete sequence of a human genome.” Science, 2022. https://doi.org/10.1126/science.abj6987 ↩︎
National Center for Biotechnology Information. “The Human Genome.” NCBI Bookshelf. https://www.ncbi.nlm.nih.gov/books/NBK595930/ ↩︎
Wang Z., Gerstein M., Snyder M. “RNA-Seq: a revolutionary tool for transcriptomics.” Nature Reviews Genetics, 2009. https://doi.org/10.1038/nrg2484 ↩︎
GTEx Consortium. “The GTEx Consortium atlas of genetic regulatory effects across human tissues.” Science, 2020. https://doi.org/10.1126/science.aaz1776 ↩︎
Human Cell Atlas. “Mapping every cell type in the human body.” https://www.humancellatlas.org/ ↩︎
Buenrostro J.D. et al. “Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin.” Nature Methods, 2013. https://doi.org/10.1038/nmeth.2688. The method uses a hyperactive Tn5 transposase, which preferentially cuts and tags DNA in regions of open chromatin; those fragments are then sequenced. ↩︎
Calo E., Wysocka J. “Modification of enhancer chromatin: what, how, and why?” Molecular Cell, 2013. https://doi.org/10.1016/j.molcel.2013.01.038. In practice, H3K4me1 is associated with primed or poised enhancers, while H3K27ac is more often associated with actively used enhancers. ↩︎
Huisken J., Stainier D.Y.R. “Selective plane illumination microscopy techniques in developmental biology.” Development, 2009. https://doi.org/10.1242/dev.022426 ↩︎
Li Lei. “From Genomics to AI Biology (Series 1).” 2026. http://lilei4409.blogspot.com/2026/06/from-genomics-to-ai-biology-series-1.html ↩︎
Avsec Ž. et al. “Effective gene expression prediction from sequence by integrating long-range interactions.” Nature Methods, 2021. https://doi.org/10.1038/s41592-021-01252-x ↩︎
Linder J. et al. “Predicting RNA-seq coverage from DNA sequence as a unifying model of gene regulation.” Nature Genetics, 2025. https://doi.org/10.1038/s41588-024-02053-6 ↩︎
Avsec Ž., Latysheva N., Cheng J. et al. “Advancing regulatory variant effect prediction with AlphaGenome.” Nature, 649, 1206–1218 (2026), published January 28, 2026. https://doi.org/10.1038/s41586-025-10014-0 ↩︎
Google DeepMind. “AlphaGenome: AI for better understanding the genome.” 2025, updated January 2026. https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome/ ↩︎
Nature’s presentation of the AlphaGenome study notes that about 98% of observed human genetic variation lies in regions that do not code for proteins. Nature Research Briefing, “Genomics: AlphaGenome predicts the impact of DNA variations,” Nature, 2026. https://www.natureasia.com/en/info/press-releases/detail/9222 ↩︎
MobiDeep is an AI meta-score for prioritizing non-coding variants in whole-genome sequencing, developed by the MoBiDiC group — Montpellier BioInformatics for Clinical Diagnosis — at the Molecular and Genomic Medicine platform, Montpellier University Hospital. It is accessible through the MobiDetails platform: https://mobidetails.chu-montpellier.fr. Group repositories: https://github.com/mobidic ↩︎
Jaganathan K. et al. “Predicting splicing from primary sequence with deep learning.” Cell, 2019. https://doi.org/10.1016/j.cell.2018.12.015 ↩︎
Cheng J. et al. “Accurate proteome-wide missense variant effect prediction with AlphaMissense.” Science, 2023. https://doi.org/10.1126/science.adg7492 ↩︎
NVIDIA. “NVIDIA BioNeMo Agent Toolkit.” https://github.com/NVIDIA-BioNeMo/bionemo-agent-toolkit ↩︎
Anthropic. “Claude Science, an AI workbench for scientists, is now available.” June 30, 2026. https://www.anthropic.com/news/claude-science-ai-workbench ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎
Polanyi M. The Tacit Dimension. University of Chicago Press, 1966. ↩︎