Voynich Interim Report
The Voynich Manuscript as a Rule-Governed System: What Our Statistical Research Has Established So Far
Short version: we have not deciphered the Voynich Manuscript, but we have found strong, repeatable evidence that its text is organized at several levels at once: paragraph openings, line-internal information flow, word formation, positional word classes, section-specific registers, Currier A/B variation, and long- versus short-distance reuse. The safest current conclusion is not “we know the language,” but “simple randomness and a single uniform generation rule are inadequate descriptions.”
8AM, 2AM, 4ODAM, SC8G, and TC8G are labels in the FSG transcription system. They are not translations, pronunciations, or claims that the manuscript literally contains Latin letters and Arabic numerals. They are machine-readable names for recurring glyph sequences.
1. Scope and method
Our main textual source is the IVTFF 2.0 corpus in the FSG transcription convention. Depending on how uncertain readings, markup, and section filters are handled, individual analyses contain approximately 32,750 to 33,330 tokens from about 202 folios. Because those parsing choices vary slightly between experiments, this report uses “about 33,000 tokens” for the corpus as a whole and gives exact counts only when they belong to a specific test.
The project combines several kinds of analysis:
- word frequency, vocabulary diversity, Zipf curves, and entropy;
- paragraph- and line-position statistics;
- prefix, root, and suffix co-selection;
- line-aware bigrams, longer windows, and slot sequences;
- Currier A/B and manuscript-section comparisons;
- image-feature and text-feature correlations on matched folios;
- copy-distance and mutation-distance tests; and
- forward simulations evaluated both in-sample and on contiguous held-out page blocks.
The distinction between description and decipherment is fundamental. A token can behave like an opener, a medial hub, or a line-final form without our knowing what it means. Labels such as “topic,” “auxiliary,” “main,” “case,” or “sentence” are therefore useful working names for statistical roles, not translations.
2. The strongest findings
2.1 The text is constrained, but not by one simple rule
The corpus combines high lexical repetition with substantial character-level diversity. In one full-corpus parsing there were 33,227 tokens and 6,719 types, with a type–token ratio of 0.202, character entropy of 3.83 bits, word entropy of 10.18 bits, and conditional bigram entropy of 2.60 bits. About two thirds of the vocabulary consisted of words occurring only once.
This combination matters. The text is not a short list of fixed phrases copied again and again, yet it is also much more formulaic than unconstrained prose. Section-level frequency curves fit Zipf-like distributions well, but their slopes differ, showing that the manuscript is not statistically homogeneous. These results rule out simple independent random-noise models. They do not, by themselves, prove ordinary natural language, meaningful content, or any particular cipher.
2.2 Paragraph openings are marked by a strong P-family signal
Words beginning with the transcriptional P- family are dramatically concentrated at paragraph openings. In the principal test, P-family rates were 38.8% at paragraph-initial position, 9.4% at ordinary line-initial position, and only 1.0% in mid-line position. The paragraph-initial versus mid-line odds ratio was 62.8, with a chi-square statistic of 2,371.9. Currier A and Currier B both showed rates close to 40%, so the effect crosses the main textual varieties.
Depending on filtering, 128–134 word forms were observed only as paragraph openers, while no comparable set was exclusive to paragraph endings. This asymmetry suggests a conventional opening mechanism rather than arbitrary lineation. The safest interpretation is that the P-family acts as a paragraph-opening or discourse-framing class. Calling it a specific conjunction, article, or grammatical particle would go beyond the evidence.
2.3 The line, not the paragraph, is the clearest compositional unit
An early analysis attributed a strong entropy decline to positions inside paragraphs. Later work separated physical lines from paragraphs and revised that conclusion. The robust effect is at the line level: vocabulary is most diverse near the beginning of a line and becomes progressively more predictable toward its end.
Across six tested sections, the correlation between word position and entropy ranged from approximately r = −0.65 to −0.91. The effect was strongest in the cosmological and biological samples and weakest, though still negative, in the long text section. This correction also explains why the earlier paragraph-level result appeared weak in long prose-like passages: some “paragraphs” span nearly an entire folio and contain many lines.
The suffix family -AK is the strongest known line-final signal. Its enrichment at line end ranged from about 2.6-fold to 61-fold across the tested sections, with the largest values in the M, stars, and long-text sections. A later segmentation produced 610 -AK-terminated spans. These are best called candidate sentence-like units, not proven sentences. After an -AK ending, the next line was 1.33 times more likely to begin with an -AM form, and an immediate -AK-to--AK restart was absent in that analysis.
2.4 Word formation follows selection constraints
Prefix and suffix choices are statistically dependent. A corpus-wide co-selection test gave chi-square = 1,549.7 with Cramér’s V = 0.082. The effect is modest in size but extremely unlikely under independent choice. Strongly enriched combinations included TO- with -AM, SC- with -ZG, and TC- with -ZG. Approximately 2,462 recurring stems were identified, and 86.7% of tokens carried a recognized prefix or suffix under that analysis.
The endings -AM, -AE, -AR, and -AK form recurring paradigmatic families across at least 85 roots. The largest measured family was 4OD-, with 462 -AM, 189 -AE, 144 -AR, and 27 -AK forms. This is good evidence for structured morphological alternation. Earlier drafts called these “nominative,” “genitive,” and “ablative” endings; that semantic identification is not established. The current defensible claim is only that these endings have different distributions and recur productively across roots.
2.5 Positional word classes exist, but their use varies by section
Frequent words separate into positional groups. Forms such as 8AM and 2AM are unusually common at line start. Forms such as SC8G, TC8G, TOE, and OE favor early-medial positions, often around position 1. Forms ending in -AK favor the end.
This initially suggested a manuscript-wide verb-second, or V2, grammar. A section-by-section check forced an important revision. The position-1 rate for the SC8G/TC8G class was about 10.0% in the biological section, 5.4% in the long-text section, 3.1% in stars, and 1.6% in herbal material, but it was absent in the tested cosmological and pharmaceutical subsets. Position-1 specialization is real, but it is a register-dependent option, not a demonstrated universal V2 grammar.
The same caution applies to proposed syntactic labels. In Currier B, line-aware adjacency produced SC8G → TC8G nine times and the reverse order once. Within a ten-token rightward window, the counts were 63 versus 30. Words from the 4O- family frequently intervene. This supports a directional frame such as SC-class → modifier slot → TC-class. Calling the two classes “auxiliary” and “main verb” remains a working analogy.
The OE family also divides into positional subtypes rather than behaving as one interchangeable particle. In the current model, 2OE is relatively start-oriented, 4OE is more bridge-like around central sequences, and some OE compounds favor terminal positions. This functional differentiation is stronger than any proposed dictionary meaning.
2.6 Currier A and Currier B are different statistical regimes
Currier A and B differ along several independent dimensions. In one integrated parsing, Currier A contained 10,546 tokens and Currier B 22,330. Average word length was 4.04 in A and 4.41 in B. Conditional bigram entropy was 2.71 bits in A and 2.43 bits in B. Their high-frequency profiles, line-position preferences, and favored morphological families also differ.
The most economical interpretation is that A and B are different generation modes or registers, not merely random spelling variation. However, B-like words also occur inside A folios in positionally restricted ways, especially away from line boundaries. That pattern is compatible with embedded technical vocabulary or local mode switching. It does not tell us whether the historical cause was dialect, scribal practice, genre, cipher setting, or some combination.
2.7 Sections have distinct statistical profiles
Section differences are not limited to subject-specific vocabulary. The herbal, biological, stars, pharmaceutical, cosmological, and long-text parts differ in vocabulary diversity, conditional entropy, average word length, opener and closer distributions, and positional syntax. For example, biological text has relatively strong local sequence constraints, whereas herbal text shows greater local variety. Long text behaves like extended prose at the paragraph scale, while some M-section material is more compatible with catalogue- or recipe-like cycling.
These observations support a multi-register manuscript. They do not identify the real-world subject of any token, and they do not prove that the conventional illustration-based section names correspond exactly to linguistic genres.
3. Long- and short-distance reuse in Currier B
A later phase asked whether recurrent forms are copied or regenerated from different memory ranges. The analysis compared line-start hubs (2AM and 8AM) with internal hubs (4ODAM, AM, OE, SC8G, and TC8G) in Currier B. For each occurrence, a likely source was sought in the preceding 80 tokens using normalized edit distance.
| Measure | Line-start hubs | Internal hubs |
|---|---|---|
| Exact reuse rate | 0.683 | 0.800 |
| Median distance to previous exact form | 48 tokens | 28 tokens |
| Mean edit distance for non-exact reuse | 0.326 | 0.278 |
The differences in exact reuse, exact-return distance, and non-exact mutation size were significant under permutation tests. The interpretation is a two-layer reset pattern: line-start hubs are reused less literally, reach farther back for exact matches, and tolerate larger mutations; internal hubs repeat more locally and conservatively. This is an observational nearest-source heuristic, not proof of literal copying by a scribe or algorithm.
8AM and 2AM also behave differently. In Currier B, 8AM had an exact-reuse rate of 0.777 and a median exact-return lag of 39 tokens; 2AM had an exact-reuse rate of 0.552 and a median lag of 64. This motivated a source-kernel model in which some returns depend on more distant prior occurrences.
4. What the simulation tests taught us
The first source-kernel model substantially improved the fit to Currier B hub statistics. Its in-sample score was 0.087 versus 0.165 for the earlier two-layer model, a relative improvement of 47.2%. It also reproduced the observed median return lag of 2AM closely: about 68 tokens in simulation versus 64 in the corpus.
The next question was harder: could more elaborate models generalize to unseen contiguous page blocks? Most did not. The table below summarizes separate experiments; lower score is better, but absolute scores should not be compared across rows because run counts and tuning grids differed.
| Extension tested | Held-out result against the original source kernel | Current verdict |
|---|---|---|
| Far-return plus local-burst mixture | 0.214 vs 0.202; 0 wins in 5 folds | Better in-sample, worse out-of-sample |
| Two-state hidden Markov model | 0.250 vs 0.218; 2 wins in 5 folds | Captures real heterogeneity, but not robustly predictive |
| Contextual semi-Markov model | 0.205 vs 0.185; 0 wins in 5 folds | Short burst memory did not generalize |
| Page-conditioned lag model | 0.216 vs 0.199; 3 wins in 5 folds | Mixed folds; worse mean score |
| Selective activation gate | 0.208 vs 0.187; 2 wins in 5 folds | Over-activates sparse or quiet regions |
| Three-state silent/sparse/active regime | Strict CV: 0.208 vs 0.187; 1 win in 5 folds | Useful description, unvalidated generator |
This negative result is scientifically valuable. 2AM clearly has quiet and burst-like periods: the fitted two-state model separated a low-rate state with a median return lag near 30 from a higher-rate state with a median lag near 5. Yet a state can be descriptively real without yielding a better forward generator. The present safe claim is:
Currier B contains state-dependent and page-local behavior, but the original long-distance source kernel remains the strongest validated single backbone. Current mixture, hidden-state, and threshold-gated extensions either overfit or use an incomplete representation of the changing state.
5. Image–text coupling: promising, but secondary
A smaller image experiment measured edge density, spread, skeleton length, branch rate, circles, concentric forms, and large connected regions on 36 sampled manuscript pages. After correcting a mistaken astronomical anchor and revising two page labels, image-based section classification improved. In the strictest excised set, three-class one-versus-rest AUC reached 0.925.
On the corrected matched-folio set, image edge density correlated positively with word entropy (r = 0.761) and the rate of an “auxiliary-like” slot class (r = 0.760), and negatively with type–token ratio (r = −0.747) and the “main-like” slot rate (r = −0.468). A previously reported spread-versus-TTR correlation weakened to p = 0.054 after the label correction and should not be treated as established.
These results suggest that visual complexity and textual organization may be coupled. They are not yet a semantic key. The image sample is small, some page-to-folio mappings were interpolated rather than visually confirmed, and classification performance can be sensitive to label decisions.
6. Claims revised or withdrawn
Several early interpretations became too strong when later tests changed the unit of analysis or introduced held-out validation. The current report therefore withdraws or narrows the following claims:
- “The entropy gradient is primarily paragraph-internal.” Revised: the strongest and most general gradient is within lines; paragraph-level behavior varies by section.
- “
-AKis principally a paragraph closer.” Revised: it is primarily a line-final suffix family and a candidate sentence-like closer. - “The manuscript has universal V2 word order.” Revised: position-1 specialization is strong in some sections and absent in others.
- “
-AM/-AE/-ARprove nominative, genitive, and ablative case.” Withdrawn as a semantic claim. A productive suffix paradigm is supported; its meanings are unknown. - “The structure proves an Indo-European or Germanic language.” Not supported. Structural analogies are not genealogical evidence.
- “A two- or three-state generator solves Currier B.” Not supported by blocked holdout. The state structure is descriptive, not yet predictively stable.
- The original astronomical page anchor. Falsified by visual audit; analyses using that mapping were corrected or treated as provisional.
7. Present working model
The findings are best summarized as a hierarchy of constraints:
- Document and section level: different sections use different statistical registers.
- Currier level: A and B are distinct but partly interacting generation modes.
- Paragraph level: the P-family strongly marks openings and frames larger discourse units.
- Line level: information diversity declines from the beginning toward the end;
-AKis a strong terminal signal. - Slot level: start-oriented, early-medial, central, and final word classes have directional co-occurrence rules.
- Morphological level: recurring roots combine selectively with prefix and suffix families.
- Memory level: line-start hubs depend on longer, more strongly mutated returns than internal hubs.
- State level: Currier B contains quiet, sparse, and active local regimes, but their predictive causes are not yet fully modeled.
This hierarchy can arise in natural language, an engineered language, a cipher with stateful rules, a structured notation system, or a hybrid process. The statistics narrow the space of explanations, but they do not select one historical interpretation on their own.
8. What should be tested next
- repeat every major positional result in independent transcription systems, with explicit glyph correspondence rules;
- use nested or strictly blocked cross-validation so that state definitions and thresholds are learned without access to the test pages;
- predict local
2AMregimes from richer page and line context instead of fixed density thresholds; - test whether one slot grammar transfers between sections or whether each section requires a different transition system;
- visually confirm uncertain page-to-folio mappings before further image–text inference;
- separate scribal, folio, section, and Currier effects in one hierarchical model; and
- seek external semantic grounding only after a structural prediction succeeds on unseen material.
9. Glossary
- IVTFF
- A tagged text format used to store Voynich transcriptions together with folio, line, section, and other annotations.
- FSG transcription
- The machine-readable glyph-naming convention used in the principal corpus for this project. Its tokens are labels, not translations.
- Folio
- A manuscript leaf. Its front and back are conventionally identified as recto and verso.
- Token / type
- A token is one occurrence of a word form; a type is one distinct word form.
- Type–token ratio (TTR)
- The number of distinct forms divided by the total number of forms. Higher values indicate greater observed vocabulary diversity, but the measure is sensitive to sample size.
- Entropy
- A measure of uncertainty or diversity. Higher word entropy means a broader, less predictable distribution of word forms.
- Conditional bigram entropy
- The uncertainty of the next symbol or token when the preceding one is known. Lower values indicate stronger local sequence constraints.
- Zipf distribution
- A common rank–frequency pattern in which a few forms are frequent and many are rare. A Zipf-like curve alone does not prove natural language.
- Hapax legomenon
- A form that occurs exactly once in the analyzed corpus.
- Prefix, root, suffix, and paradigm
- Working divisions of a transcribed form into an initial part, a recurring central family, and a final part. A paradigm is a set of related forms created by systematic alternation. These are distributional units here, not proven pronunciations or meanings.
- Slot grammar
- A model in which classes of forms prefer certain positions and transitions, such as opener → early-medial class → modifier class → terminal class.
- Currier A and Currier B
- Two long-recognized statistical varieties of Voynich text. The names describe recurring differences; they do not by themselves specify language, dialect, author, or cipher.
- Hub
- A frequent form whose occurrence organizes many following contexts or recurrence patterns. In this project,
8AMand2AMare important line-start hubs. - Lag or copy distance
- The number of tokens between a form and a possible earlier source or exact match.
- Normalized edit distance
- The number of insertions, deletions, or substitutions needed to transform one form into another, adjusted for form length.
- Source kernel
- A probabilistic rule that changes the chance of generating a form according to how far away a related earlier occurrence lies.
- Hidden Markov model (HMM)
- A model in which observed events are generated by unobserved states that change probabilistically, such as “quiet” and “burst” regimes.
- Blocked holdout / cross-validation
- A test in which contiguous page groups are withheld during model fitting and used only for evaluation. It is stricter than measuring fit on the same pages used to choose the model.
- AUC
- Area under the receiver operating characteristic curve, a classification measure ranging from 0.5 for chance-level ranking to 1.0 for perfect ranking.
- p-value
- The probability, under a stated null model, of observing a result at least as extreme as the one measured. It is not the probability that a hypothesis is true.