UNIFIED STATE SYSTEM
UNIFIED STATE
LOVE & FREEDOM — ALWAYS
UNIFIED STATE · Blogs

Language, images & the next reader

Blog · Language, images & the next reader

As Unified State Language begins to carry whole books, stable addresses and recoverable images become parts of the same practical question: what does the next reader need?

A sentence is a generous first test for a communication system. It is easy to inspect, quick to encode and small enough to keep in mind. A book asks more. It brings repeated words, unusual spellings, punctuation, long passages and a source that nobody should have to compare by eye after transmission.

Our latest work brings that larger scale into Unified State Language. The lexical composer can process book-sized text, both image instruments have larger grids and automatic fitting, and Carrier and the lexical base now share enforced colour ownership. Each improvement helps the reader know what arrived and how to interpret its references.

A book supplies a useful test

The release check used a 4,455,950-byte plain-text edition of the King James Bible. It produced 1,386,973 lexical-message tokens and 1 image frame at a 1,872 × 1,872 data grid. The saved 3,792-pixel-square PNG was reopened and read back; the recovered source matched the original byte for byte through its SHA-256 fingerprint.

The example matters because it exercises a real long text. It also exposes an earlier practical restriction: the composer stopped at one million message tokens. Removing that ceiling required more careful processing. A further check repeated the book three times: 13,367,850 source bytes and 4,160,917 message tokens also completed the image recovery path exactly.

The updated tools bound both source text and the resulting message packet at 128 MiB. Message metadata can make a near-limit source exceed the packet budget, and the browser still needs working memory. Those practical limits make the supported scope explicit; the book checks establish the measured examples within it.

This measured result applies to that source and those encoding settings. Another book can produce a different token count, compressed size and number of images. The instruments show those figures for the actual message being composed.

A spelling and a concept have different jobs

English Lexical Base 1 still contains 354,983 source-attested forms. This upgrade preserves the frozen collection and its published identifiers. Long documents become usable through changes to the message tools, without a bulk rewrite of the dictionary.

A form such as bank has an address in the lexical collection. Its appearance in a sentence does not, by itself, select a financial institution or a river’s edge. Carrier gives us a separate place to develop concepts, definitions and relations with evidence and review.

The composer can therefore use lexical references for recognised forms, retain other passages as literal text and attach a Carrier reference where the sender deliberately selects one. The original punctuation and spacing still travel with the message. More addressable vocabulary creates useful structure; explicit references make the intended connections inspectable.

An address should remain dependable

That structure needs a common allocation rule. Before the global registry, different layers could describe their references precisely while still using colour addresses independently. As the language grows, accidental overlap becomes harder to notice and more expensive to repair.

The shared registry now reserves addresses across Carrier and lexical releases. A new Carrier entry cannot take a colour already owned by an English form or another concept. A future vocabulary import must respect the same reservations. Dropping an entry retains its claim, and promotion from working to checked preserves its allocation identity.

The migration preserved every English form address. It corrected three duplicate Carrier colours and kept the original dictionary histories. Older messages still need their pinned context: an old swatch that once had competing uses cannot explain its intended reference on its own.

Address ownership and knowledge quality remain distinct responsibilities. The registry can establish which record owns an address. A definition, translation or claimed relationship needs its own grounds for acceptance.

Choose the route the message needs

The two instruments now explain their roles in one line. Visual carries exact text or file bytes directly through CVP1. Lexicon first represents text with English form addresses and optional Carrier concepts, then uses the same image transport.

For a document that should arrive as the same file, the direct Visual route is straightforward. For text whose vocabulary and selected conceptual connections should be available to the receiver, the lexical route adds that structure. Both provide a defined path back from image data to source.

Lexicon’s file chooser accepts UTF-8 text. It preserves line endings and any UTF-8 byte-order mark while the imported source remains unedited. Binary files belong in Visual. This small distinction helps avoid quietly changing a file while presenting it as exact recovery.

Let the square fit the message

Both instruments offer data grids of 64, 128, 256, 1,024, 2,048 and 4,096 cells per side. Automatic fitting considers the final stream after compression, including the metadata needed to read it. It chooses the smallest supported square that can hold that stream; a larger stream continues across maximum-size frames.

The selected grid describes the data area. CVP1 also needs its landmark border, header, checks and error correction. A maximum 4,096-cell data grid therefore produces a complete 4,120-pixel square when each tile occupies one pixel. The interface reports the actual dimensions and occupied capacity instead of treating source length as image size.

A large grid gives more room, while a sequence gives the message room to continue. Exact original PNGs provide the reference material for recovery. Resizing, blur and camera capture introduce separate questions that require measured tests.

Seeing every possible colour

Visual also offers a different kind of image: the full namespace atlas. Its native 4,096 × 4,096 pixels enumerate all 16,777,216 RGB24 addresses, each once. Exact lookup connects a colour to its coordinate, and the original PNG preserves the complete grid.

This view includes both assigned and available addresses. The global directory supplies ownership information. A fitted preview helps us navigate the whole square; inspecting the original coordinates reveals individual values.

CVP1 has another task. It repeatedly uses sixteen calibrated transport colours to carry encoded bytes, with landmarks and correction information. The atlas shows the available address space; a message image carries a particular communication through an agreed format.

What another reader can check

Image sequences can travel with source text, message data, hashes and the dictionary context required by explicit references. These paired exports make repeatable evaluation possible and provide material for future model-training experiments. Any claim about learned recognition or understanding will need its own results.

For now, the practical invitation is simple: choose a text, decide which route suits it, export the images and read them back. A reader should be able to recover the words, inspect the references and see the limits of the method. Carrying a book makes those expectations concrete.