Boyce Data Science / special topic in omics
Reading the Dark Genome
A staged guide to the large share of human DNA that does not code for protein: what the term covers, what the literature shows, who is working on it, which therapies target it, how findings are interpreted, and how the data are stored, analyzed, and governed.
Start here
The guide shows one stage at a time. Use the bar above or the buttons at the foot of each stage to move through it, in order or by jumping to the stage you need.
Choose the description closest to you and each stage will open with a short note written for that reader, along with a suggested reading order. You can change the choice whenever you like, and the full content stays available to everyone.
Status of this guide
This guide is under construction. It is an educational overview, written for a general professional and community readership, and it has not been peer reviewed and is not medical advice. If you see an error, a missing paper or program, or a better way to explain something, please write to danielle@boycedatascience.com. Statements about companies, trials, and affiliations are a snapshot taken in September 2026 from the public sources listed on the last stage, and each one includes a way to check the current status. Several entries rest on company press releases, which describe results in the sponsor's own words.
Save your place, your notes, and your shortlist
Nothing you enter is stored in the browser or sent anywhere. To keep your reader choice, notes, data source shortlist, and draft statement, download a save file and load it again later.
What the term covers
"Dark genome" is an informal label, and authors use it for different things. Reading any paper or press release on the topic starts with working out which meaning the author has in mind.
The non-coding genome
The most common use: all DNA that does not code for protein. A 2026 review in the European Heart Journal describes roughly 1 to 2 percent of the genome as protein-coding and lists the remainder as regulatory elements, transposable and repetitive sequences, structural features such as telomeres and centromeres, pseudogenes, introns, intergenic regions, and non-coding RNA genes. Some writers call the same territory genomic "dark matter".
Regions that short-read sequencing cannot resolve
A technical use: stretches of the genome where standard short reads cannot be placed with confidence, because the sequence is repeated elsewhere or because coverage drops out. Ebbert and colleagues (2019) called these dark and camouflaged regions and showed that they include parts of genes already linked to disease.
Understudied protein-coding genes
A drug discovery use: proteins that are coded in the genome but have attracted little research. The Illuminating the Druggable Genome (IDG) program of the National Institutes of Health (NIH), begun in 2014, labels the least studied proteins "Tdark" in its Pharos portal. These genes are protein-coding, so this meaning overlaps little with the first.
The dark proteome
A newer use: small proteins and peptides translated from open reading frames that reference annotation does not list as genes, often inside long non-coding RNAs or untranslated regions. Ribosome profiling finds thousands of these, and a community effort led through GENCODE is cataloging them.
This guide treats the first meaning as the main subject and brings in the other meanings where they change how data are collected or interpreted.
The main components
How much of the dark genome is functional: the debate and the estimates
In 2012 the Encyclopedia of DNA Elements (ENCODE) consortium reported that about 80 percent of the genome shows some reproducible biochemical activity, such as transcription or protein binding, in some cell type. Critics, including Graur and colleagues (2013), answered that biochemical activity is a much looser standard than function shaped by natural selection. Comparative genomics gives a lower figure: Rands and colleagues (2014) estimated that about 8 percent of the human genome is under purifying selection, most of it non-coding.
The two figures measure different things and both remain in use. A cautious reading is that a sizable minority of non-coding DNA has sequence-dependent function, a larger share is biochemically active without clear evidence of function, and the boundary between the two is an open research question. "Junk DNA", a phrase from Susumu Ohno in 1972, is a historical term and not a finding.
History of the idea
The science behind the dark genome developed over decades, from the discovery of mobile DNA to complete genome assemblies and sequence models that score non-coding variants. The timeline can be narrowed to ideas, technology, clinical findings, or therapies.
Literature review
The literature is organized here by theme, with a short account of what each body of work shows and where it is contested. The reading list at the end of the stage can be filtered by theme, and every entry links to the paper or to a PubMed search that finds it.
Reading list: reviews, landmark papers, and guidance, by theme
Keeping a literature review current
The field moves quickly, so a dated search string is more durable than a fixed list. A starting PubMed query is "dark genome"[tiab] OR "non-coding genome"[tiab] OR "noncoding variant*"[tiab] OR "transposable element*"[tiab], narrowed with your disease term. Record the date, the query, and the count each time you run it, and keep preprints (bioRxiv, medRxiv) in a separate list because their conclusions may change before publication. Run the "dark genome" search in PubMed.
People and groups
The researchers, consortia, companies, communicators, and meetings listed here are entry points for following the field. The list is a starting set drawn from the literature in this guide and is not a ranking; many productive groups are not named.
Affiliations change. Those shown are as the sources in this guide gave them, and a name search in PubMed or on the institution's site will confirm where someone works now.
Consortia, programs, and meetings where the field gathers
Books and periodicals for general readers and for staying current
Therapeutic approaches
Most medicines act on proteins. Work on the dark genome adds targets that are RNA molecules, regulatory DNA, or the products of reactivated repeat elements, and each kind of target pairs with a different kind of drug.
The examples include investigational programs described by their sponsors. The next stage states, for each one, whether the information comes from a regulator, a peer-reviewed journal, a trial registry, or the company.
Reading claims about a dark genome therapy
Several questions help when a program is described in a press release or a talk. Which element is the target, and is the evidence that it drives disease genetic, correlational, or from a model system? Is the reported result a biomarker change or a change in how people feel and function? Was the comparison randomized and blinded, and was the outcome named in advance? Has the result appeared in a peer-reviewed paper or on a trial registry, or only in a company statement? A favorable biomarker result in a small early trial is a reason for a larger trial, and many such results have not held up at the next step.
Drug development timeline
The entries below run from approved medicines whose target lies in non-coding sequence, through programs now in trials, to earlier programs that stopped. Together they show that the dark genome has already reached the clinic in a few places and that most of the field is still early.
The basis of the information on each card
Every card names what its information rests on, followed by a plain statement of what that means, and both come before the description of the program. Much of what is known about drugs in development comes from the companies developing them. A press release or company website is the sponsor's own account: it has not been checked by independent scientists or by a regulator, and it is written partly for investors. It belongs in this guide as a record of what a company says it is doing, and it is not scientific evidence that a treatment works.
An investigational drug is available only through a clinical trial. Designations such as Fast Track, Orphan Drug, Rare Pediatric Disease, and Breakthrough Therapy change how a company works with the FDA and are not findings that a drug is effective. Decisions about care belong with a person's own clinicians.
What each basis means, from regulatory approval to a company website
Status is a snapshot from September 2026. Each card has a link that runs the search on ClinicalTrials.gov, and the last stage lists the press releases and papers behind each entry.
Filter by stage of development
Filter by the basis of the information
The usual path from a non-coding target to an approved medicine
- Target finding. Genetics, functional genomics, or expression studies point to an element. Human genetic support tends to predict later success better than expression differences alone.
- Target validation. Perturbing the element in cells or animals changes a disease-relevant readout. Many long non-coding RNAs are poorly conserved across species, so animal models can be limited and human cell systems do more of the work.
- Drug design. The modality is chosen to fit the target: an antisense oligonucleotide (ASO) for an RNA, an editor for a DNA element, a small molecule for an enzyme such as a reverse transcriptase.
- Investigational New Drug (IND) application and phase 1. Safety, dosing, and drug levels, often in healthy volunteers.
- Phase 2. Early evidence of effect, frequently through biomarkers such as neurofilament light chain (NfL) in neurology.
- Phase 3 and review. Larger trials with clinical outcomes, then review by the Food and Drug Administration (FDA) or another regulator. Designations such as Orphan Drug, Rare Pediatric Disease, and Fast Track change the process and incentives but are not judgments that a drug works.
- Individualized therapies. For a variant found in one person or a few, a custom ASO may be developed under an expanded access IND, as with milasen. This path depends on having the genome sequence, a variant mechanism an ASO can address, and an institution able to run it.
Clinical interpretation
Interpreting a non-coding variant draws on the same framework used for coding variants, with different evidence. The first practical question is whether the test that was ordered could have seen the region at all.
What each kind of test can see
The discovery of ReNU syndrome shows the consequence of test scope. In 2024, independent groups reported that variants in RNU4-2, a small non-coding RNA gene of the spliceosome, explain an estimated 0.4 percent of neurodevelopmental disorders. The gene makes no protein, so it is absent from most exome designs, and the diagnosis had been out of reach for families whose testing stopped at the exome. Related findings in RNU2-2 and RNU5B-1 followed.
Adapting the classification framework
Diagnostic laboratories classify variants with the 2015 standards from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP). Ellingford and colleagues (2022) published recommendations for applying those standards outside coding sequence. In outline, they advise restricting analysis to regions with an established link to the gene and phenotype in question, defining candidate regulatory elements from experimental data in a relevant tissue, using computational predictors built for non-coding sequence, and weighting functional assays that measure the predicted mechanism, such as RNA sequencing to show a splicing change.
Kinds of evidence used for a non-coding variant, with their usual sources
Established examples of disease and drug response linked to non-coding sequence
Reanalysis over time
A non-coding variant of uncertain significance (VUS) may be reclassified as knowledge grows, and a genome stored today can be searched again when a new gene such as RNU4-2 is reported. ACMG has published points to consider on reevaluation and reanalysis, and the American Society of Human Genetics (ASHG) has a position statement on recontacting research participants. Laboratories and clinics differ in whether reanalysis happens automatically, on request, or not at all, so the policy in force is a fair question to ask of any provider.
Computational scores, including those from recent sequence models, are one line of evidence. Published recommendations treat them as supporting evidence that does not classify a variant without segregation, population, phenotype, or functional data alongside it.
Patients and advocates
For families and advocacy groups, the dark genome bears on two practical things: whether past genetic testing looked in the right places, and how a community can prepare its samples and data so that new discoveries reach its members.
In plain language
Genes that make proteins take up a very small part of our DNA. The rest includes the switches that turn genes on and off in each tissue, genes that make working RNA molecules instead of proteins, and long stretches of repeated sequence left by ancient viruses and mobile DNA. Changes in these regions can cause disease, and many common genetic tests do not read them.
When a test came back negative
A negative result describes what that test examined. Gene panels and exomes read mainly protein-coding sequence, so a negative exome leaves most non-coding causes unexamined. A genome sequence reads far more, although some repeated regions still need long-read methods. A genetic counselor can say which test was done and what it could and could not detect.
Questions to bring to a genetics appointment
How an advocacy group can prepare
Talking about early-stage therapies
Families often hear about dark genome programs through company announcements. Groups can help by describing the stage of each program plainly, separating biomarker results from results about daily function, naming the source, and giving the date the information was checked. The drug development timeline in this guide labels every program with the basis of its information, such as a company press release or a peer-reviewed publication, and the same labels can be used in a community update. The questions on the therapeutic approaches stage can serve as a template.
Ethics
Most ethical questions in genomics apply here, and the dark genome sharpens several of them because far more sequence is read than can be interpreted, and because interpretation will keep changing for years after a sample is taken.
Organizational statements
Professional societies and standards bodies have published guidance that touches on non-coding findings, and an advocacy group, clinic, or registry may want a short statement of its own. This stage lists existing guidance first and then helps you draft a statement.
Existing guidance and position statements that bear on the dark genome
No society had issued a statement devoted to the dark genome as a whole when this list was compiled in September 2026; the documents above address parts of it. Search the society sites for newer documents before citing this list.
Draft a statement for your organization
Fill in what applies, choose the positions your organization is ready to hold, and generate a draft. The text is a starting point for your board, medical advisors, and community to revise, and every position is optional.
Your draft will appear here after you choose positions and select Generate the draft.
Data sources
Public resources for the dark genome fall into a few families: reference sequence and annotation, functional genomics, catalogs of non-coding RNA, repeats and translated open reading frames, population and clinical variation, tools that score variants, target knowledge bases, and cohorts with genome sequence. Add entries to a shortlist and it will be included in your take-home file.
Access levels
Reference and functional genomics resources are generally open. Individual-level human genomes are held under controlled access, through the database of Genotypes and Phenotypes (dbGaP), the NHGRI Analysis, Visualization, and Informatics Lab-space (AnVIL), the UK Biobank Research Analysis Platform, the All of Us Researcher Workbench, or the Genomics England Research Environment. Expect a data use agreement, institutional sign-off, and analysis inside a secure environment that does not permit individual-level data to be downloaded.
Data forms
Dark genome data arrive as raw reads, alignments, variant calls, signal tracks, interval lists, contact matrices, graphs, and annotation files, and each form has a place where non-coding information is commonly lost.
Identifiers, nomenclature, and standards for exchanging non-coding findings
Genome builds and coordinates
A non-coding variant is identified by its coordinates, so the reference build has to travel with every record. Human data are in circulation on GRCh37 (hg19), GRCh38 (hg38), and the telomere-to-telomere assembly T2T-CHM13, and converting between them can fail or shift positions in the repetitive regions that this field studies. Store the build, the annotation release (for example the GENCODE version), and the transcript used for naming, and avoid merging files from different builds without an explicit conversion step that logs what could not be converted.
Approaches to analysis
Analytic methods are grouped here by the question being asked. Tool names are examples in common use and not endorsements; most questions have several reasonable tools, and the choice often follows what your collaborators already run.
Recurring pitfalls in dark genome analysis
Statistical considerations for analysts from a clinical data background
The search space is the main difference from coding analysis. A genome holds millions of non-coding variants per person, so tests that work across coding genes lose power unless the analysis is restricted in advance to defined elements, such as promoters and enhancers active in the relevant tissue. Rare variant burden tests need a principled unit to aggregate over, and the choice of unit should be fixed before looking at outcomes. Fine-mapping and colocalization give probabilities and not certainties, and they inherit the ancestry makeup of the reference panels. Deep learning scores are predictions from sequence trained on a limited set of cell types, so performance reported on benchmark sets may not transfer to a tissue the model never saw.
Informatics
Bioinformatics, clinical informatics, and research informatics each handle a different part of this work: the first produces and annotates the calls, the second decides how results reach the health record and the clinician, and the third joins genomic findings to phenotype data under governance that holds up over time.
Bioinformatics
- Keep aligned reads in CRAM, or keep the raw reads, so that the genome can be realigned to a newer reference and searched again; a file of coding variant calls alone cannot be reanalyzed for non-coding causes.
- Run callers for the variant classes that short-read pipelines skip by default: structural variants, short tandem repeat expansions, and mobile element insertions.
- Annotate against regulatory and RNA gene catalogs as well as protein-coding transcripts, and record the versions of every reference and annotation file.
- Use workflow managers (Nextflow, Snakemake, WDL) and containers so that a rerun years later reproduces the first run.
- Plan storage and compute: a 30x genome is tens of gigabytes as CRAM, and long-read and single-cell assays are larger.
Clinical informatics
- Many genetic results reach the electronic health record (EHR) as scanned or PDF reports, which cannot be queried. Structured results using Health Level Seven (HL7) Fast Healthcare Interoperability Resources (FHIR) Genomics and Logical Observation Identifiers Names and Codes (LOINC) make a variant findable later.
- Record the test scope with the result, because "negative" from a panel and "negative" from a genome mean different things for non-coding causes.
- Decide how reclassified variants and reanalysis findings return to the ordering clinician and who owns that workflow.
- Pharmacogenomic (PGx) decision support already acts on non-coding variants, such as promoter and intronic alleles in the Clinical Pharmacogenetics Implementation Consortium (CPIC) guidelines, and offers a working model for alerts and result storage.
Research and registry informatics
- The Observational Medical Outcomes Partnership (OMOP) Common Data Model has no native home for genome-scale variant data. Common practice is to keep genomic files in a purpose-built store, link them to OMOP person identifiers, and bring selected findings into OMOP as measurements or observations; the Observational Health Data Sciences and Informatics (OHDSI) community has worked on genomic vocabulary and extensions.
- Global Alliance for Genomics and Health (GA4GH) Phenopackets package phenotype terms from the Human Phenotype Ontology (HPO) with variant interpretations, which suits case-level exchange for rare disease.
- Keep a person, specimen, assay run, and file identifier chain so that an RNA result can be traced to the same person's DNA result.
- Write consent and governance that permit reanalysis, recontact, and controlled-access sharing, since the value of a stored genome grows as the non-coding catalog grows.
A reference pipeline from sample to interpreted non-coding finding
Related Boyce Data Science tools
For the wider setting, see From Clinical Data to Omics for analysts with a clinical data background who are new to omics projects, and Following the data for tracing each data type through a multimodal registry.
Glossary
Every abbreviation used in the guide is spelled out here, with a plain-language meaning.
Sources and take-home
Download a take-home file with your reader note, shortlist, notes, and draft statement, then review the web sites and papers this guide was built on.
Anything you want to keep from your reading: open questions, names to follow, data sets to request.
Web sites and papers used
Links open in a new tab. Status claims about therapeutic programs come from the press releases and registry entries listed under "Therapeutic programs" and reflect what those pages said in September 2026.