The Space Reviewin association with SpaceNews
 


 
Tsiolkovsky
Konstantin Tsiolkovsky‘s archive was digitized years ago, but not in a machine-readable or searchable form. (credit: Konstantin E. Tsiolkovsky State Museum of the History of Cosmonautics)

The letters Tsiolkovsky asked for

A 1934 fireball, 970 pages of replies, and the first machine-readable edition of the father of astronautics’ personal archive


On May 14, 1934, a fireball passed over the Moscow region. Five weeks later, on June 21, the newspaper Izvestia printed a short note by Konstantin Tsiolkovsky, titled “Who saw the bolide?”, asking eyewitnesses to write to him in Kaluga.

As of August 2026, the entire fond—2,019 archival files, 51,008 scanned pages—exists for the first time as a machine-readable corpus.

They did. His personal archive preserves 221 files, or 970 pages, of replies: a physics teacher at an agricultural technical school, a livestock specialist, a foreman at a chemical plant, a head of a technical control department, a magazine editor, a village council. The responses spanned weeks: 195 letters in June, 20 more in July. Tsiolkovsky annotated the letters and began an article, “On the bolide of May 14, 1934,” which he never finished.

For 90 years those letters have been sitting in fond 555 of the Archive of the Russian Academy of Sciences in Moscow, Tsiolkovsky’s personal fond. They were scanned and posted online years ago. But “scanned” turned out to be a long way from “readable,” and the story of closing that gap is what this article is about.

As of August 2026, the entire fond—2,019 archival files, 51,008 scanned pages—exists for the first time as a machine-readable corpus: catalogued, dated, page-classified, and fully transcribed, released into the public domain under CC0 with a permanent DOI and a public search interface. The bolide correspondence is one small, human-scaled corner of it, and for the first time anyone can search it and read it.

One caveat up front: the transcription was produced by machine and has not been verified by a human on a single page. Its accuracy has been measured—that measurement, I would argue, is the most interesting part of the project—and the corpus is a finding aid, not a citable edition. Every file links back to the archive’s own scans, which remain the source of record. Nothing previously unknown was “discovered.” Rather, what changed is that the material became usable.

Why “digitized” did not mean “accessible”

The archive’s portal serves the scans, but working with it as a data source turns up obstacles that only appear when you try to use it.

The identifier in a page’s web address does not match the archival file number: id=300 corresponds to file 297 of the first inventory, because 31 files carry lettered numbers (145a, 077b, 585a, and others) that occupy positions in the running numeration without shifting the archival one. Large scans, above roughly 800 kilobytes, are served incomplete, cut off around 130 kilobytes, and can only be retrieved with byte-range requests; a browser silently saves the truncated file. And, in at least four cases, the “variant number” in a file’s description contradicts its own dating: a “second variant” dated September 1, 1920, precedes a “first variant” dated October 19 of the same year.

None of this is a complaint about the archive, which did the essential work of scanning and posting 51,008 pages. It is instead an illustration of the distance between putting images online and making a fond queryable. Closing that distance took a catalog of all 2,019 files (published July 30, 2026, with a DOI), dating for 1,969 of them (155 marked as tentative), a per-page classification—34,903 manuscript pages, 14,585 typescript, 1,520 covers and annotations—and, finally, transcription of every page.

The measurement problem, and a trick the typewriter era left us

Machine transcription of handwriting is easy to produce and hard to trust. To measure its accuracy you need ground truth: the same text keyed in by a human. For an archival fond, no such ground truth exists; if it did, machine reading would be unnecessary. The existing workarounds lean on a model’s internal confidence signals.

Dating 1,969 files from the portal’s catalog cards makes it possible, apparently for the first time, to look at the whole fond on a time axis.

Personal archives of the typewriter era offer a way out. They often preserve one text twice: the autograph manuscript and a typescript copy made from it, filed together. You can read both with the same pipeline and compare the transcriptions. The text is the same and the pipeline is the same; what differs is only the legibility of the page. The disagreement between the two readings is, almost entirely, a measure of how hard the handwriting is.

Fond 555 contains 1,759 such autograph-typescript pairs across 224 files. The two readings agree, at the median, on 37% of the words. The median longest verbatim run shared between them is 10 words, and in 43% of pairs no shared run exceeds that. Those numbers are the honest calibration of what “machine-read manuscript” means here.

Is the trick itself trustworthy? It can be checked against real ground truth in the minority of cases where one exists: files whose text also survives in a printed edition. There the paired-reading estimate is unbiased to within a percentage point, and it ranks pages by difficulty the same way the truth does: a rank correlation of 0.67 across all 55 available pairs, rising to 0.92 on the 32 pairs where the printed edition is demonstrably the same text, and 0.97 on the 17 most certain.

Against printed editions directly, the earlier pipeline measured 98.1% character accuracy on typescript and 81.1% on manuscript (91.7% and 73.7% at word level). On the completed corpus the picture splits and cannot be averaged: where an archival file contains the same redaction that was printed, character accuracy is 92.3%, and where the file holds a draft or working materials toward an article, agreement with the edition drops to 24%. But that is the distance between a draft and its published form, not reading error.

Reporting the negative results

A project like this generates findings that would be tempting to leave out. They belong in the article as much as the successes do.

Comparing authorial redactions is impossible at this reading quality. The fond preserves “The Space Ship” in two variants, exactly the material for textual comparison. The variants agree with each other on 19% of words, or less than two machine readings of the same page agree with each other. Authorial revision cannot be distinguished from reading error, so that line of work was closed rather than published. The prohibition is built into the tool: the software refuses to display differences that fail the measured threshold.

The page classifier passed its test and was still wrong. It separates manuscript from typescript by the spread of ink-stroke lengths, was checked against 19 hand-labeled pages, and got all 19 right. On the 4,675 pages where an independent signal existed to check against, its accuracy is 80%: carbon copies and faded typescript read as manuscript.

Two pages cannot be read by any model at all: a German typescript review from 1927 of one of Tsiolkovsky’s brochures trips a built-in filter against verbatim reproduction of known printed text. They were read by a separate route and are flagged in the corpus.

What the dating shows

Dating 1,969 files from the portal’s catalog cards makes it possible, apparently for the first time, to look at the whole fond on a time axis. Eighty-six percent of the files fall in the last 18 years of Tsiolkovsky’s life. The 1930s account for 58% of the files but only 40% of the pages. In his final years his work does not stop, but it shortens from extended treatises to brief notes. The 1920s contribute 557 files and the 1910s 145. Everything before 1890 amounts to single files, and the post-1935 dates belong to materials about Tsiolkovsky rather than by him.

Against that curve, the bolide correspondence of 1934 stands out. A 77-year-old who had moved to short notes still turned a newspaper into a distributed observation instrument, and several hundred readers—teachers, technicians, a village council—answered within weeks. The unfinished article the letters were meant to feed is in the fond too.

Transcribing history for $39

The remainder of the fond—41,212 pages—was transcribed in one overnight batch run for approximately $39. The model was chosen by measurement rather than by price list: candidates were compared on manuscript pages against the typescript copy of the same text, with the previous pipeline as the baseline on the same pages. On typical handwriting no model clearly outran the rest; on a 48-page sample the more expensive model was better on 37 pages and worse on 11—a real but modest edge that the material, not the price, turns out to limit.

The fond that documents his life is now, for the first time, something anyone can search, with its accuracy printed on the label.

One practical finding deserves passing on. Newer models enable “reasoning” by default, and it spends the same output budget as the answer. With a 4,096-token limit, up to 3,929 tokens went to reasoning; the transcription broke off mid-page, and the stub looked like bad reading rather than a truncated response. With reasoning explicitly set to minimum, quality did not drop and output volume halved.

What this is and is not

The corpus is published only for files transcribed in full, because a partial transcription reads as continuous text with a missing middle. Uncertain readings are flagged word by word and illegible passages, authorial deletions (38,465 of them), insertions and marginalia are marked; pre-reform orthography is preserved as written. The dataset is CC0 with a permanent DOI (10.5281/zenodo.21705221, always resolving to the latest version, currently v1.4.2). The code is on GitHub. The full-fond search includes an English catalog search alongside the Russian. A preprint describing the method is on the arXiv. The scans remain in the archive: the corpus links to them rather than republishing them.

An English translation of the corpus was deliberately not made, though its cost was measured at $29 for the whole fond: translating a machine transcription stacks error on top of error and uncertainty flags vanish in translation, leaving an English text that would look more reliable than the Russian it came from. Expert verification has not been performed on any page and is not planned here: that is work for a manuscript specialist, not a machine.

The 61st Tsiolkovsky Readings opened in Kaluga on September 15, two days before what would be his 169th birthday. The fond that documents his life is now, for the first time, something anyone can search, with its accuracy printed on the label.


Note: we are now moderating comments. There will be a delay in posting comments and no guarantee that all submitted comments will be posted.

Home