The paper is available as a PDF:
Budargham, B. (2026). The EOA Program, Part 0: The Engineering of Alphabets as an Open Frontier. GitHub. https://github.com/bahaa-budargham/eoa-part0-paper
We introduce the Engineering of Alphabets (EOA), a research program built on an observation that, as far as we can tell, current theories of representation do not predict. At the center of the program is a recursive operator we call beta. It takes the spoken names of alphabet letters and maps them into numerical sequences. When we run beta on the 26 letters of English under a bijective phonetic mapping, the sequences converge to a limiting constant ratio (LCR). For the English primary mapping, that ratio is about 3.306. Under other encodings we tested, it shifts: 3.037, 3.173, or 2.689.
Here is the strange part. The 26 sequences do not stay independent. They collapse into what looks like a 16-element basis. Build word vectors from that basis, and they land on a very low-dimensional subspace. We hypothesize, on the basis of our unreleased runs, that a recursive, tree-like generative structure drives the collapse; formalization and verification are pending (Section 11). In our own runs of the operator, these phenomena are stable and vary with the encoding. Independent testing is not possible yet, because the operator is not public. We plan to open it to collaborators under the academic collaboration terms described in Section 12.
This paper is Part 0 of the EOA Program. It offers no proofs and no conclusive experiments. Its job is to pose the questions that Parts I through V will try to answer, and to place those questions where they belong: at the intersection of symbolic dynamics [4], phonetic encoding theory, low-rank representation learning, and deterministic machine learning.
All rights reserved. Access to files is granted for review or reference upon request. Reproduction, redistribution, or commercial use is prohibited without express permission.
- EOA Program, Part I (coming soon)
- EOA Program Main Page
- EOA Data