howclose.to
Intelligence & Machines · Updated July 2026Momentum · steady

How close are we to practical DNA data storage?Can DNA become a real computer archive?

or, simply: Can DNA become a real computer archive?or, precisely: How close are we to practical DNA data storage?

Synthetic DNA has recovered a 200 MB random-access archive without bit errors and automated a complete write-to-read cycle for five bytes, but no public benchmark joins useful payload, throughput, retrieval and cost.DNA can hold computer files in an extraordinarily small sample. The missing step is a useful machine that writes a large archive quickly, finds one file and costs less than today's cold storage.

We are here

Writing and system integration remain the bottlenecks - A field assessment still identified cost-effective high-throughput writing, practical storage handling and integrated workflows as barriers to deployment.

01 · Where we stand

State of playWhere dna data storage stands right now

The current stage, the honest metric, and the single threshold that gates the next stage. Each threshold is a falsifiable claim with a named next test.How far up the ladder we've climbed, the honest verdict, and the one thing blocking the next step.

The five stagesMaturity ladder
Deployed
Out in the real worldDeployed at scale
Scaling
Making it cheap enough at scaleScaling toward competitive cost
Engineering← HERE
Building one that pays for itselfEngineering a system that pays back
Lab demo
Shown to work in a labDemonstrated in the laboratory
Theoretical
The idea is worked out on paperTheoretical basis established
The honest verdictVerified state

DNA can hold computer files in an extraordinarily small sample. The missing step is a useful machine that writes a large archive quickly, finds one file and costs less than today's cold storage.Synthetic DNA has recovered a 200 MB random-access archive without bit errors and automated a complete write-to-read cycle for five bytes, but no public benchmark joins useful payload, throughput, retrieval and cost.

200Best result so far35 files · no bit errors
1,000,000,000One-petabyte operational archive · editorial scale test
Blocking the next stepBlocking threshold

Write it fast enough and cheaply enoughArchive-class cost and throughput Next test: An appliance vendor must publish an audited end-to-end dollars-per-TB and bytes-per-second benchmark.

The verdict, in five rungsHow far up the ladder↓ next — the thresholds & the gap
01 · The evidence

The thresholds that gate the next stageWhat has to happen next

Each threshold is a falsifiable claim with a named next test; the gap chart shows how far today's metric sits from the goal.Each row is one thing that has to be proven — and how far today's number is from the target.

Lossless multi-file archiveGet a large set of files back✓ Achieved · Mar 2018
100%
Proven byMicrosoft and University of Washington · 200 MB
Automated end-to-end cycleMake the whole process automatic✓ Achieved · Mar 2019
100%
Proven byFive bytes recovered after an approximately 21-hour write-to-read cycle
Archive-class cost and throughputWrite it fast enough and cheaply enoughEarly
0%
Next testAn appliance vendor must publish an audited end-to-end dollars-per-TB and bytes-per-second benchmark.
Operational petabyte archiveRun a real petabyte archiveEarly
0%
Next testA deployment must publish usable payload, write and retrieval performance, integrity checks and operating cost together.
THRESHOLDS - Thresholds for DNA Data Storage.
Scale
Fully recovered research-archive payload over time, with measured values, projected values, and a goal at 1.0e+9 MB user data.1101001,00010,000100,0001.0e+61.0e+7Fully recovered research-archive payload · MB user dataYear20132018GOAL 1.0e+9 · One-petabyte operational archive…Five files · 100% accurate…35 files · no bit errors~1.0e+9 MB user data to goal
NOTE - This archive-payload ladder is a proxy because no comparable public series combines payload, write cost, throughput, random access and automation. The one-petabyte endpoint is an editorial scale test, not a forecast. The automated 2019 system is excluded from the archive curve and shown separately because it handled only five bytes.
02 · How we got here

The record behind the verdict

Major events set large; context events set small but never hidden. Everything below the TODAY rule is a schedule, not a result.

1964-20111 event1 shown

Molecular memory is imagined

Molecular memory is imagined begins with nucleic-acid memory is proposed. The result established the next question for the field.

1964
Nucleic-acid memory is proposedTheory
Mikhail Neiman discussed the possibility of using nucleic-acid macromolecules as non-biological information memory.
2012-20162 events2 shown

Digital files enter DNA

Digital files enter DNA moved the field from a 5.27-megabit book enters dna to five files return with 100% accuracy. The results narrowed the next question without closing it.

2012
A 5.27-megabit book enters DNAExperiment
A book containing more than 53,000 words, 11 images and a program was encoded in about 55,000 synthetic DNA strands. This was an encoded-payload demonstration, not the later no-bit-error archive record.
2017-20181 event0 shown

Density and random access

Density and random access begins with dna fountain demonstrates extreme density at small scale. The result established the next question for the field.

2017
DNA Fountain demonstrates extreme density at small scale
A 2.1 MB experiment reported 215 PB/g physical density. Synthesis cost about $7,000 and sequencing about $2,000, so density did not imply an affordable storage system.
2019-20242 events1 shown

Automation and molecular search

Automation and molecular search moved the field from automation completes a five-byte cycle to a molecular index searches by image similarity. The results narrowed the next question without closing it.

2019
Automation completes a five-byte cycleExperiment
The first automated end-to-end device wrote, stored and recovered HELLO. The five-byte cycle took about 21 hours and only one of 30 extractable payload reads decoded perfectly in the reported run.
2021
A molecular index searches by image similarity
A DNA-based index performed similarity search over representations of 1.6 million images, demonstrating computation near the stored molecules rather than a larger archival payload.
2025-20262 events1 shown

The integrated-system test

The integrated-system test moved the field from cas9 adds addressed and semantic retrieval to writing and system integration remain the bottlenecks. The results narrowed the next question without closing it.

2025
Cas9 adds addressed and semantic retrieval
Cas9-guided methods selectively enriched files and tested semantic image lookup. The semantic database represented 1.74 million images with 457 synthesized address sequences; it was not a 1.74-million-file archive.
2025
Writing and system integration remain the bottlenecksSetbackWe are here
A field assessment still identified cost-effective high-throughput writing, practical storage handling and integrated workflows as barriers to deployment.
2018-20181 event1 shown

Events outside the declared eras

Events outside the declared eras begins with random-access archive reaches 200 mb. The result established the next question for the field.

2018
Random-access archive reaches 200 MBExperiment
Thirty-five files totaling 200 MB were encoded in more than 13 million oligonucleotides, selectively accessed and fully recovered with no bit errors.
— end of record · 9 shown, 0 hidden —
9 events · below the TODAY rule = scheduled, not done
03 · The data behind the verdict

Why the meters read the way they do

The learning curves and comparisons that justify each threshold's percentage. Every series is measured, with the source event linked in the timeline above.

The molecules work. The archive does not-yet.The recovered archive reaches 200 MB; the 1 PB operational test is unmeasured.One dot = one fully recovered archive · payload proxy in MB · log scale · 2013–2018 · automation and density are separate demonstrationsOne dot = one fully recovered archive · MB user data · log scale · 2013–2018 · five-byte automation, molecular density and random access are separate measures; the petabyte endpoint is an editorial estimate
DNA archive payload proxy and separate automation resultThe solid archive ladder shows fully recovered research payloads only. The one-petabyte endpoint is an undated editorial scale test, not a forecast. The five-byte automated write-store-read result sits on a separate rail because it did not match the archive payload.FULLY RECOVERED RESEARCH-ARCHIVE PAYLOAD · PROXY METRIC100 kB10 MB1 GB100 GB10 TB1 PB739 kB2013 · Five files · 100% accurate recovery200 MB2018 · 35 files · no bit errors5,000,000× PAYLOAD SCALE GAP1 PBEDITORIAL SCALE TESTNO DATE · NOT A FORECASTAUTOMATION PROOF · SEPARATE AXIS5 BHELLO · ~21-hour cycle · 2019PAYLOAD RECORD ≠ AUTOMATION ≠ DENSITY ≠ AFFORDABILITY
NATURE · MICROSOFT RESEARCH · AS OF JUL 2026
04 · What it unlocks

If the remaining tests pass

Downstream capabilities, drawn dashed because they depend on results not yet in.

DNA Data StorageUltra-dense cold archivesKeep rarely read scientific and cultural records in much less physical space than today's tape libraries.Low-standby-energy preservationPreserve offline information without continuously powering disks or refreshing semiconductor memory.Search inside molecular librariesSelect or compare information in the storage medium before decoding the entire pool.
05 · Sources

Where every number comes from

9 sources — every figure on this page traces to one.