Blog · Language, meaning & visual communication
With 354,983 source-verified English forms, a separate lexical layer brings ordinary text into the visual language while leaving meaning open to careful, explicit work.
We have reached a useful point in Unified State Language: the infrastructure can now carry a substantial English vocabulary alongside the concepts already taking shape in Carrier.
The first release, USL English Lexical Base 1, makes 354,983 source forms addressable. Its standalone instrument lets people explore the vocabulary, compose exact text, attach selected Carrier concepts, and send the resulting message through colour images. It also reads those images back.
The architectural choice behind this release matters as much as its size. English becomes one vocabulary mapped onto Unified State Language. Carrier keeps its role as the place where concepts, definitions, relations and their evidence receive deliberate attention.
A word opens several doors
Consider bank. The same spelling can point toward a financial institution or the edge of a river. Giving that spelling an address is straightforward. Deciding which meaning a sentence expresses requires more context.
Likewise, run, runs, running and ran invite a relationship between forms. A reliable account would also distinguish the many senses of run: moving on foot, managing an organisation, or operating a machine.
This leads us toward four distinct entities: surface form → lexeme → sense → concept. The visible spelling is a form. A lexeme groups appropriate forms of a vocabulary item. A sense identifies a particular use. A concept can connect that use with a wider structure of meaning.
- FormObserved spelling
Available in this release - LexemeRelated forms
Future sourced mappings - SenseA particular use
Future sourced mappings - ConceptMeaning and relations
Checked or working in Carrier
Connections between these roles need explicit evidence. An imported spelling does not create the connections automatically.
This release establishes surface forms. Lexeme grouping, sense distinctions and concept mappings remain work to be sourced and reviewed. That boundary allows vocabulary to grow quickly without turning unexamined assumptions into dictionary meaning.
A vocabulary with a traceable beginning
Our baseline comes from the Moby Words II single-word collection preserved by Project Gutenberg. Its documentation records the author’s public-domain grant. We retain the original source files and notices with the release.
The measured download contains 354,983 unique, nonempty forms, one fewer than the count advertised in the source documentation. We preserve that discrepancy. A reproducible collection should describe the file it actually imported.
Every form has an identifier and a record of its source line. The release manifest pins the source and derived files through cryptographic hashes, making the import independently checkable.
Here, “source-verified” has a precise scope: the form occurs in the preserved source. It supplies no automatic definition, grammatical analysis, translation or endorsement. The collection is a documented English baseline, with the history and limitations of its source.
Giving forms stable addresses
RGB24 provides 16,777,216 possible colour addresses. This release occupies approximately 2.116% of that capacity within its lexical namespace.
The namespace is essential. A colour alone cannot tell a reader which vocabulary, release or meaning it belongs to. Our messages identify the lexical namespace, version and exact release hashes before using the form identifiers. The same RGB value elsewhere does not establish semantic equivalence.
Existing identifier-to-form bindings are intended to remain stable as this vocabulary grows. New allocations must preserve those bindings. Corrections need explicit records that retain the history, so an older message continues to point to the form its author used.
Exact words, explicit meaning
The composer combines lexical references, literal text and explicit Carrier references in one message.
Exact vocabulary matches can use lexical identifiers. Remaining text travels literally as UTF-8. That preserves punctuation, whitespace, unfamiliar words and other writing systems. Message text is not silently normalised to fit the imported vocabulary.
An author can also select a passage and attach a Carrier concept. For example, the phrase reciprocal legibility can carry a reference to its reviewed concept while retaining the words in the original sentence. A spelling match alone never creates that attachment.
Each concept reference records its registry, checked or working tier, entry and exact dictionary context. If the reader lacks that context, the original text survives and the concept remains visibly unresolved. The system also leaves the appropriateness of the attachment open to interpretation and review.
From a message to an image
The visual transport uses the separately versioned CVP1 protocol. Its landmarks, framing, integrity checks and error correction give encoded messages a defined structure.
A lexical RGB address and a transmission colour have different jobs. CVP1 uses sixteen calibrated palette colours, with four payload bits per data module. The protocol does not assume that every image pixel can reliably carry an arbitrary 24-bit vocabulary address.
Larger messages span image sequences. Frames carry the information needed to assemble the sequence correctly, even when their PNG files arrive out of order. The receiver still needs the pinned vocabulary release and the relevant context for any concept references it wants to resolve.
The standalone instrument embeds the complete first vocabulary release and a saved Carrier snapshot. It also attempts to load and verify live checked and working context, identifying which context is available.
Measuring what the images carry
The release checks recovered exact source bytes from five example messages across six PNG frames. Those examples include an explicit concept reference, Unicode and line endings, and a longer sequence assembled from reversed frame order.
In the long example, 3,500 source forms required two images; plain UTF-8 with gzip required three under the same CVP1 settings. Short messages showed the cost of release information and framing. The instrument therefore displays the actual comparison for each message.
These results establish useful transport behaviour. They do not establish universal compression, camera recovery or a model’s ability to understand an image unaided.
The work this makes possible
The next layers can connect forms to lexemes, distinguish senses and build carefully supported relationships across languages. Ukrainian, Spanish and other vocabularies can acquire their own identities while sharing concepts where the evidence supports that connection.
Image exports also pair frames with original text, typed messages, hashes and available context. Those reproducible pairs offer material for future recognition experiments and training datasets. Model training and measured recognition performance remain ahead of us.
This extends the questions explored in our earlier AI perspective and The Pixel Is Not the Meaning: how can visual communication carry useful structure while keeping its interpretation inspectable?
Explore the English Lexical Base, compose a message, attach a concept deliberately, and read the image back. Ordinary vocabulary now has a place in the system. The work of connecting it to meaning can proceed with a clearer foundation.