% !TEX program = xelatex % % Epiphany --- Text Projection (companion specification) % Companion to the Core Specification. Compile with XeLaTeX. % % This document is versioned independently of the Core Specification % (independent semver; see the Versioning note in the front matter). Its preamble % is intentionally a self-contained copy of the core specification's preamble so % the two documents build independently; factoring a shared preamble file is a % later cleanup, not a v0.1 deliverable. \documentclass[11pt,letterpaper]{report} % --------------------------------------------------------------------------- % Packages % --------------------------------------------------------------------------- \usepackage{fontspec} \usepackage{geometry} \geometry{ letterpaper, top=1.05in, bottom=1.05in, left=1.15in, right=1.15in, headheight=15pt } \usepackage[english]{babel} \usepackage{microtype} \usepackage{parskip} \usepackage{xcolor} \usepackage{hyperref} \usepackage{enumitem} \usepackage{titlesec} \usepackage{fancyhdr} \usepackage{booktabs} \usepackage{array} \usepackage{longtable} \usepackage{listings} \usepackage{amsmath} \usepackage{amssymb} \usepackage{tcolorbox} \tcbuselibrary{breakable, skins} % --------------------------------------------------------------------------- % Color palette (shared with the core specification) % --------------------------------------------------------------------------- \definecolor{epiphanyteal}{HTML}{1A4044} \definecolor{epiphanygold}{HTML}{8E6E2E} \definecolor{epiphanyink}{HTML}{1F1B16} \definecolor{epiphanyslate}{HTML}{6B6660} \definecolor{epiphanycream}{HTML}{F8F4ED} \definecolor{epiphanymist}{HTML}{ECE8E0} \definecolor{epiphanycode}{HTML}{2A2520} \definecolor{epiphanycrimson}{HTML}{7A2424} \hypersetup{ colorlinks=true, linkcolor=epiphanyteal, citecolor=epiphanyteal, urlcolor=epiphanygold, pdftitle={Epiphany --- Text Projection}, pdfauthor={The Epiphany Project}, pdfsubject={Text Projection companion for the Epiphany music notation platform}, pdfkeywords={music notation, operations, CRDT, reduction, serialization}, bookmarksnumbered=true, bookmarksopen=true } % --------------------------------------------------------------------------- % Typography (shared with the core specification) % --------------------------------------------------------------------------- \setmainfont{TeX Gyre Pagella}[Numbers={OldStyle, Proportional}, Ligatures={TeX, Common}] \setsansfont{TeX Gyre Heros}[Scale=0.94, Ligatures={TeX, Common}] % No TeX ligatures in the mono font: `tlig` maps " to a right curly quote and -- % to an en dash. This document's grammar quotes terminals with U+0022 and spells % escapes as backslash sequences, so a substituted glyph would misstate the syntax. \setmonofont{TeX Gyre Cursor}[Scale=0.88] \newfontfamily\titlefont{TeX Gyre Pagella}[Numbers={OldStyle}, Ligatures={TeX, Common}] \newcommand{\tablenums}[1]{{\addfontfeatures{Numbers={Lining,Tabular}}#1}} \newcommand{\sectionsc}[1]{{\addfontfeatures{Letters=SmallCaps}#1}} % --------------------------------------------------------------------------- % Section styling (shared with the core specification) % --------------------------------------------------------------------------- \titleformat{\chapter}[display] {\normalfont\filright} {\raggedright\color{epiphanygold}\fontsize{14pt}{16pt}\selectfont \scshape Chapter\ \thechapter} {16pt} {\raggedright\color{epiphanyteal}\fontsize{32pt}{36pt}\selectfont\bfseries} [\vspace{4pt}{\color{epiphanygold}\rule{2in}{0.6pt}}] \titlespacing*{\chapter}{0pt}{-20pt}{30pt} \titleformat{\section} {\normalfont\Large\bfseries\color{epiphanyteal}} {\color{epiphanygold}\thesection}{1em}{} \titleformat{\subsection} {\normalfont\large\bfseries\color{epiphanyteal}} {\color{epiphanygold}\thesubsection}{1em}{} \titleformat{\subsubsection} {\normalfont\normalsize\bfseries\color{epiphanyink}} {\thesubsubsection}{1em}{} % --------------------------------------------------------------------------- % Headers and footers (shared with the core specification) % --------------------------------------------------------------------------- \pagestyle{fancy} \fancyhf{} \renewcommand{\headrulewidth}{0pt} \renewcommand{\footrulewidth}{0pt} \fancyhead[L]{\small\scshape\color{epiphanyslate}Epiphany --- Text Projection} \fancyhead[R]{\small\itshape\color{epiphanyslate}\leftmark} \fancyfoot[C]{\small\color{epiphanyslate}\thepage} \renewcommand{\headrule}{ \color{epiphanygold!50}\hrule width\headwidth height 0.4pt \vspace{1pt} \color{epiphanygold!30}\hrule width\headwidth height 0.2pt } % --------------------------------------------------------------------------- % Code listing style (shared with the core specification) % --------------------------------------------------------------------------- \lstdefinelanguage{Rust}{ keywords={fn,let,mut,pub,struct,enum,impl,trait,for,in,if,else,match,return, use,mod,crate,self,Self,as,where,move,async,await,const,static, ref,type,unsafe,extern,dyn,box,break,continue,loop,while}, keywordstyle=\color{epiphanyteal}\bfseries, ndkeywords={i8,i16,i32,i64,i128,u8,u16,u32,u64,u128,f32,f64,bool,char,str, String,Vec,Option,Result,Box,Rc,Arc,HashMap,BTreeMap, NonZeroU16,NonZeroU32,NonZeroU64,Duration,Timestamp}, ndkeywordstyle=\color{epiphanygold}\bfseries, sensitive=true, comment=[l]{//}, morecomment=[s]{/*}{*/}, commentstyle=\color{epiphanyslate}\itshape, stringstyle=\color{epiphanycrimson}, morestring=[b]", morestring=[b]' } \lstset{ basicstyle=\ttfamily\small\color{epiphanycode}, backgroundcolor=\color{epiphanycream}, frame=leftline, rulecolor=\color{epiphanygold!60}, framesep=8pt, framerule=1.5pt, xleftmargin=10pt, xrightmargin=4pt, breaklines=true, showstringspaces=false, numberstyle=\tiny\color{epiphanyslate}, numbersep=10pt, captionpos=b, aboveskip=10pt, belowskip=10pt, language=Rust } % --------------------------------------------------------------------------- % Custom environments (shared with the core specification) % --------------------------------------------------------------------------- \newtcolorbox{openquestion}[1][]{ enhanced, breakable, colback=epiphanymist, colframe=epiphanycrimson, fonttitle=\bfseries\color{white}, title={\scshape\hspace{2pt}Open Question}, coltitle=white, colbacktitle=epiphanycrimson, arc=1pt, boxrule=0pt, leftrule=2pt, left=10pt, right=10pt, top=8pt, bottom=8pt, attach boxed title to top left={xshift=0pt, yshift=0pt}, boxed title style={arc=0pt, sharp corners, boxrule=0pt, left=6pt, right=8pt, top=2pt, bottom=2pt}, #1 } \newtcolorbox{rationale}[1][]{ enhanced, breakable, colback=epiphanymist, colframe=epiphanyteal, fonttitle=\bfseries\color{white}, title={\scshape\hspace{2pt}Rationale}, coltitle=white, colbacktitle=epiphanyteal, arc=1pt, boxrule=0pt, leftrule=2pt, left=10pt, right=10pt, top=8pt, bottom=8pt, attach boxed title to top left={xshift=0pt, yshift=0pt}, boxed title style={arc=0pt, sharp corners, boxrule=0pt, left=6pt, right=8pt, top=2pt, bottom=2pt}, #1 } % Numbered within chapter (this document has chapters); see core_spec.tex's % requirement box for why a plain counter + `code=` step is used instead of % tcolorbox's own "auto counter, number within=..." keys. \newcounter{requirement}[chapter] \renewcommand{\therequirement}{\thechapter.\arabic{requirement}} \newtcolorbox{requirement}[1][]{ enhanced, breakable, colback=white, colframe=epiphanygold, fonttitle=\bfseries\color{white}, code={\refstepcounter{requirement}}, title={\scshape\hspace{2pt}Requirement~\therequirement}, coltitle=white, colbacktitle=epiphanygold, arc=1pt, boxrule=0pt, leftrule=2pt, left=10pt, right=10pt, top=8pt, bottom=8pt, attach boxed title to top left={xshift=0pt, yshift=0pt}, boxed title style={arc=0pt, sharp corners, boxrule=0pt, left=6pt, right=8pt, top=2pt, bottom=2pt}, #1 } \newtcolorbox{nongoal}[1][]{ enhanced, breakable, colback=epiphanymist, colframe=epiphanyslate, fonttitle=\bfseries\color{white}, title={\scshape\hspace{2pt}Non-Goal}, coltitle=white, colbacktitle=epiphanyslate, arc=1pt, boxrule=0pt, leftrule=2pt, left=10pt, right=10pt, top=8pt, bottom=8pt, attach boxed title to top left={xshift=0pt, yshift=0pt}, boxed title style={arc=0pt, sharp corners, boxrule=0pt, left=6pt, right=8pt, top=2pt, bottom=2pt}, #1 } \newcommand{\MUST}{\textbf{MUST}} \newcommand{\MUSTNOT}{\textbf{MUST}\nobreak\ \textbf{NOT}} \newcommand{\SHOULD}{\textbf{SHOULD}} \newcommand{\SHOULDNOT}{\textbf{SHOULD}\nobreak\ \textbf{NOT}} \newcommand{\MAY}{\textbf{MAY}} \setlist[itemize]{topsep=2pt, itemsep=3pt, parsep=0pt} \setlist[enumerate]{topsep=2pt, itemsep=3pt, parsep=0pt} \setlist[description]{topsep=2pt, itemsep=5pt, parsep=0pt} \AtBeginDocument{\color{epiphanyink}} % --------------------------------------------------------------------------- % Document % --------------------------------------------------------------------------- \begin{document} \begin{titlepage} \thispagestyle{empty} \centering \vspace*{2.2in} {\color{epiphanygold}\rule{3in}{0.8pt}}\\[18pt] {\titlefont\fontsize{34pt}{38pt}\selectfont\color{epiphanyteal}\bfseries Epiphany}\\[10pt] {\Large\scshape\color{epiphanyslate}Text Projection}\\[6pt] {\large\itshape\color{epiphanyslate}A companion to the Core Specification}\\[14pt] {\color{epiphanygold}\rule{3in}{0.8pt}}\\[24pt] {\normalsize\color{epiphanyink}Version 0.14.0 --- Canonical bases leave the companion: the container epoch carries what text cannot}\\[4pt] {\small\color{epiphanyslate}Normative for the text form it defines} \vfill \end{titlepage} \tableofcontents % =========================================================================== \chapter{About This Companion} \label{ch:about} The \emph{Text Projection} is a companion to the Epiphany Core Specification. It fulfils the delegation the core specification makes in Chapter~8, \sectionsc{Text Projection} (\texttt{sec:format:textproj}), which declares that ``the format admits a deterministic projection to a canonical s-expression text form'' and that ``the text projection is normative'' --- while leaving the form itself unwritten. The Binary Format companion likewise excludes it: ``the canonical s-expression form --- that is the \emph{Text Projection} companion's''. This document supplies that form. \section{What This Document Covers} \begin{itemize} \item The canonical text syntax: atoms, byte strings, text, and the layout that makes the projection deterministic. \item What is projected, and what is deliberately not. \item The projection and parse requirements, including the bidirectional round-trip the core specification demands. \end{itemize} It does \emph{not} cover the binary encoding of anything --- that is the Binary Format companion's --- nor the semantics of any operation, which is the Operation Catalog's. The projection is a \emph{re-presentation} of the canonical document, never a second definition of it. Wherever the two could disagree, the binary form is normative and the projection is wrong. \section{The Subject of the Projection} The projection's subject is the \textbf{canonical document}, not the file. Per the core specification's Chapter~8 \sectionsc{Text Projection}, the projection \MUSTNOT{} be required to preserve chunk offsets, compression choices, cache chunks (operation indexes, layout caches, integrity indexes), garbage bytes from prior commits, or superblock generation numbers, slot assignments, and CRCs. Consequently two bundles that differ only in physical layout project to the same text, and a text re-serializes to \emph{a} bundle rather than to \emph{the} bundle it came from. That is the intent, not a limitation: the physical file is an encoding of the document, and the projection is of the document. A bundle \MAY{} cache its own projection in a \texttt{TextProjection} chunk named by \texttt{Manifest.text\_projection\_root}. That chunk is a \textbf{non-canonical accelerator} (core specification Chapter~8, \sectionsc{Schema Versioning}): a reader need not understand it, and a writer preserves it verbatim or discards it. A cached projection that disagrees with the operations it claims to project is \emph{stale}, not authoritative. % =========================================================================== \chapter{The Canonical Text Form} \label{ch:form} \section{Encoding and Character Set} \begin{requirement} \label{req:textproj:charset} A text projection \MUST{} be UTF-8, with no byte-order mark. Every text field it carries \MUST{} be in Unicode NFC, matching the canonical-text rule of the core specification's Appendix~D, \sectionsc{Text and Unicode}. A parser \MUST{} reject non-NFC text rather than normalize it. \end{requirement} \section{Atoms} An \emph{atom} is a symbol, an integer, a byte string, or a text string. \textbf{Symbols} are lowercase ASCII words, possibly hyphenated: \texttt{envelope}, \texttt{insert-event}, \texttt{strict-inverse}. They name constructors and enumeration cases. A symbol is never quoted. \textbf{Integers} are written in base ten, with a leading \texttt{-} for negative values, no leading zeros, and no leading \texttt{+}. Zero is \texttt{0}, never \texttt{-0}. \textbf{Byte strings} carry every identifier, hash, and opaque payload. \begin{requirement} \label{req:textproj:hex} A byte string \MUST{} be written as \texttt{\#x} followed by an even number of \textbf{lowercase} hexadecimal digits, one pair per byte, in the order the bytes appear in the canonical binary form. The empty byte string is \texttt{\#x}. A parser \MUST{} reject uppercase digits, an odd digit count, and any separator within the digits. \end{requirement} \begin{rationale} One rule for every byte string. Hexadecimal has no alphabet variant and no padding to canonicalize, it is greppable, and a corrupted character is locally obvious. Identifiers are 16 or 32 bytes, so its expansion costs nothing where it is read; only an inlined snapshot pays. The core specification permits ``base64 \emph{or another canonical text form}''; a second encoding would buy a quarter of the bytes of the one body nobody reads, at the price of pinning an alphabet, a padding rule, and a line-wrapping rule, and of choosing which encoding applies where. Ratified at 0.1.0. \end{rationale} \textbf{Text strings} are double-quoted. Inside a string, \texttt{\textbackslash{}"} denotes a quotation mark, \texttt{\textbackslash{}\textbackslash{}} a backslash, \texttt{\textbackslash{}n} a line feed, and \texttt{\textbackslash{}t} a tab; no other escape exists. \begin{requirement} \label{req:textproj:string-escapes} A text string \MUST{} escape exactly the characters that require it: the quotation mark, the backslash, U+000A, and U+0009. Every other character \MUST{} appear literally. A parser \MUST{} reject an escape sequence outside this set, and \MUST{} reject a literal character that the writer was required to escape. \end{requirement} \begin{rationale} ``Escape exactly'' rather than ``escape at least'': the text is canonical, so two spellings of one string cannot both be valid. This is the same injectivity the binary form rests on (Binary Format, \texttt{req:binfmt:decode-vectors}), stated for text. \end{rationale} \section{Layout} \begin{requirement} \label{req:textproj:envelope-per-line} A projection is a sequence of lines separated by a single U+000A, with a final U+000A and no other trailing whitespace. Each line is one complete s-expression. Tokens within a line are separated by exactly one space; there is no other whitespace, and no indentation. Each operation envelope \MUST{} occupy exactly one line. \end{requirement} \begin{rationale} The core specification's stated use case is that ``merge conflicts surface at the operation-envelope level, which is the meaningful level for collaborative editing''. One envelope per line makes a line-based three-way merge conflict \emph{exactly} an envelope conflict --- never a conflict inside an envelope, which could otherwise produce a syntactically valid operation that neither side wrote. It also disposes of indentation: there is no whitespace to canonicalize, so the ``identical semantics project to identical text'' requirement below has nothing to hide in. The lines are long. Readability is a \emph{tooling} concern, and a pretty-printer is free to reformat for display; what it must not do is write the reformatted text back and call it a projection. Ratified at 0.1.0. \end{rationale} % =========================================================================== \chapter{What Is Projected} \label{ch:content} \section{Derive, or Carry --- Never Both} \label{sec:content:derive-or-carry} The manifest's references are \emph{physical}. A \texttt{ChunkRef} carries an offset, a compressed length, and a compression algorithm; a \texttt{BlobRef} carries the same. Those are exactly the things the projection \MUSTNOT{} preserve. But a reference also carries an \emph{identity} --- a chunk id, a content hash, a blob id --- and every one of those is a function of the content. One rule resolves both, and it is the rule Requirement~\ref{req:textproj:reduced-state-derived} already applies to reduced state. \begin{requirement} \label{req:textproj:derive-or-carry} A projection \MUST{} carry exactly what the document does not determine, and \MUSTNOT{} carry anything it does. \begin{itemize} \item \textbf{Physical attributes} --- a chunk's or blob's \texttt{offset}, \texttt{compressed\_length}, and \texttt{compression}, and a chunk's \texttt{uncompressed\_length} --- \MUSTNOT{} appear. A serializer chooses them freely. \item \textbf{Derivable identities} --- \texttt{ChunkId}, \texttt{ContentHash}, \texttt{BlobId} --- \MUSTNOT{} appear. They are re-derived from the content by the derivations the Binary Format companion pins (\sectionsc{Content Hashing}, \sectionsc{Domain-Separated Preimages}). \item \textbf{Non-derivable identities} \MUST{} appear. In schema major~0 there is exactly one: \texttt{SnapshotId}, which the Binary Format companion declares opaque, with readers forbidden from deriving it (\texttt{req:binfmt:snapshot-id-opaque}). \item \textbf{Content and semantic attributes} \MUST{} appear: a chunk's \texttt{kind} and \texttt{schema\_version} and its uncompressed payload; a blob's media type, declared maximum uncompressed length if any, and its payload. \end{itemize} A parser \MUST{} reject a projection carrying a value this requirement forbids. \end{requirement} \begin{rationale} Carrying a derivable identity would reproduce, at the level of a chunk, the defect Requirement~\ref{req:textproj:reduced-state-derived} rules out at the level of the document: two sources of truth for one fact, with nothing to stop them disagreeing. Carrying a physical attribute would make two encodings of one document project to two texts, breaking Requirement~\ref{req:textproj:canonical-text}. \texttt{SnapshotId} is the sole exception, and it is an exception for a stated reason rather than an oversight: v0 has no snapshot producer, so the identity has nothing to derive \emph{from}, and the Binary Format companion accordingly pins it as sixteen opaque bytes that a reader \MUSTNOT{} attempt to verify. A projection must carry what it cannot recompute. \end{rationale} \section{Document Structure} \label{sec:content:structure} A projection is, in order: \begin{enumerate} \item a \texttt{(text-projection )} header line, naming the version of \emph{this companion} the text conforms to; \item a \texttt{(document \#x (schema ))} line --- the manifest's carried \texttt{SchemaVersion} is projected alongside the document id, verbatim and never re-derived (Requirement~\ref{req:textproj:manifest-schema-carried}) --- and a \texttt{(lineage \#x)} line if the manifest declares one; \item zero or more \texttt{(profile ...)} lines, in canonical order; \item zero or more \texttt{(extension ...)} lines, in canonical order; \item at most one \texttt{(canonical-base ...)} line; \item zero or more \texttt{(blob ...)} lines, in canonical order; \item zero or more \texttt{(envelope ...)} lines, in canonical operation order (core specification Appendix~D). \end{enumerate} \begin{requirement} \label{req:textproj:manifest-schema-carried} The \texttt{document} line \MUST{} carry the manifest's aggregate \texttt{SchemaVersion} (core specification Chapter~8, \sectionsc{Schema Versioning}) as its second field. A binary-to-text projector \MUST{} copy this value verbatim from the bundle superblock, and a text-to-binary serializer \MUST{} use the value declared by the document line verbatim as the bundle's manifest schema version. This companion \MUSTNOT{} derive or validate the value from \texttt{ExtensionDeclaration.edit\_barriers}; those bytes are opaque at this layer. A document author editing opaque barrier bytes is responsible for updating the field to match. \end{requirement} \begin{requirement} \label{req:textproj:header-version} A parser implementing this companion \MUST{} accept exactly one header version: \texttt{(0 14 0)}, the version of the companion it implements. It \MUST{} reject any other version at line one. Multi-version acceptance and text migrate-on-read are deferred in the same posture as op-payload migrate-on-read: support belongs in an explicit, version-keyed migration path when a real consumer requires it. The current parser \MUSTNOT{} speculate by accepting another version. \end{requirement} Every sequence is written in the normative order its binary counterpart uses. The projection introduces no ordering of its own. \section{Canonical Blobs} \label{sec:content:blobs} \begin{requirement} \label{req:textproj:canonical-blobs} A blob referenced by a canonical operation or by canonical reduced state is itself canonical (core specification Chapter~8, \sectionsc{Canonical and Non-Canonical Roots}). Every such blob \MUST{} be projected as a \texttt{(blob ...)} line carrying its media type, its declared maximum uncompressed length if it declares one, and its uncompressed payload. Its \texttt{BlobId}, content hash, offset, lengths, and compression are re-derived (Requirement~\ref{req:textproj:derive-or-carry}). The blob lines are ordered and de-duplicated by their projected form (Requirement~\ref{req:textproj:derived-ordering}), not by the binary order, which reads the offset. A blob referenced only by acceleration structures is non-canonical and \MUSTNOT{} be projected. Parser handling of unreferenced blob lines is specified by Requirement~\ref{req:textproj:reject-unreferenced-blobs}. \end{requirement} \begin{rationale} An embedded image, font, or audio recording that a canonical operation references is part of the document. Omitting it would make the projection lossy for exactly the documents most in need of archival --- and lossy \emph{silently}, since the operations that reference the blob would still be there, pointing at a blob id the text no longer contains. This requirement was absent from version 0.1.0 of this companion, and from the core specification's own list of what the projection preserves; both are corrected. \end{rationale} \begin{requirement} \label{req:textproj:reject-unreferenced-blobs} A parser \MUST{} reject every \texttt{(blob ...)} line whose blob is unreferenced by canonical state (Requirement~\ref{req:textproj:canonical-blobs}). At companion version~0.14.0, neither a canonical operation nor canonical reduced state can carry a \texttt{BlobId}; canonical state therefore cannot reference a blob, and a parser \MUST{} reject every \texttt{(blob ...)} line. \end{requirement} \begin{rationale} A blob-bearing text at this version is necessarily non-canonical. Accepting one stages a blob into the bundle that the next projection silently drops, causing data loss and falsifying $\textrm{project}(\textrm{serialize}(\textrm{parse}(T))) = T$ for that text. Forward compatibility belongs to header-version gating (Requirement~\ref{req:textproj:header-version}), not to leniency here. \end{rationale} \section{Profile Declarations} \label{sec:content:profiles} A \texttt{(profile ...)} line carries the profile's identity, its semantic version, and its constraints. \begin{requirement} \label{req:textproj:profile-id} A profile identity is a symbol for each closed-vocabulary profile (\texttt{full}, \texttt{read-only}, \texttt{lite}), and \texttt{(custom \#x)} for \texttt{ProfileId::Custom}, whose registry id is sixteen bytes. \end{requirement} \begin{rationale} Version 0.1.0 required a symbol for every profile, which made a custom profile unrepresentable and its claim to preserve ``all profile declarations'' false. \end{rationale} \section{Extension Declarations} \label{sec:content:extensions} \begin{requirement} \label{req:textproj:extension-declaration} An \texttt{(extension ...)} line \MUST{} carry every field of the declaration: its identity, its semantic version, whether it is required, its affected object kinds, its edit barriers, and its preserved chunk roots. The affected object kinds and the edit barriers are opaque to the bundle and are projected as byte strings, verbatim. Fields appear in the ratified declaration order (core specification Chapter~8, \sectionsc{Extension Declarations}). Each preserved chunk root is projected as its \texttt{kind}, its \texttt{schema\_version}, and its uncompressed payload (Requirement~\ref{req:textproj:derive-or-carry}), never as a \texttt{ChunkRef}: a \texttt{ChunkRef} is a physical reference, and the projection has no file to point into. The chunk list is ordered and de-duplicated by projected form (Requirement~\ref{req:textproj:derived-ordering}), because \texttt{ChunkRef}'s binary order breaks ties on the offset. \end{requirement} \begin{rationale} \texttt{affected\_object\_kinds} and \texttt{edit\_barriers} have ratified structured shapes (\texttt{ObjectKind}, \texttt{EditBarrier}) \emph{and} canonical byte encodings, and the bundle stores them opaquely: it preserves them across reads and writes without interpreting them. At 0.3.0 the projection does the same, carrying their canonical bytes, on the principle that the projection interprets nothing the bundle does not. This is a deferral, not a conclusion. A later revision \MAY{} project them structurally under Requirement~\ref{req:textproj:value-projection}; because their canonical bytes are unchanged by that, doing so does not change the document, only its text --- and that is a version-gated change to this companion, not to the format. \end{rationale} \begin{rationale} Version 0.1.0 carried only identity, the required flag, barriers, and a list of chunks, dropping the semantic version and the affected object kinds outright and leaving the chunks' representation undefined. An extension's declaration governs how the canonical document is \emph{interpreted}; a projection that loses part of it does not determine the document. \end{rationale} \section{Ordering What the Binary Form Ordered Physically} \label{sec:content:derived-ordering} ``Keep the binary form's order'' is available only where the binary order is a function of data the projection preserves. Two sequences fail that test. \begin{requirement} \label{req:textproj:derived-ordering} Where a sequence's binary order depends on attributes Requirement~\ref{req:textproj:derive-or-carry} erases, the projection \MUST{} order its elements by their \textbf{projected form}, ascending, comparing the UTF-8 bytes of the rendered element; and \MUST{} emit at most one element per distinct projected form. In schema major~0 this applies to exactly two sequences: \begin{itemize} \item the \texttt{(blob ...)} lines, whose binary counterpart \texttt{blob\_roots} is sorted by the full \texttt{BlobRef} encoding --- which contains the offset, the compressed length, and the compression algorithm; and \item an extension's preserved chunk roots, sorted in binary by \texttt{ChunkRef}'s order, whose key is the kind, then the content hash, then the \textbf{offset}, with the compressed length, the uncompressed length, and the compression as further tie-breakers. \end{itemize} Every other projected sequence keeps the binary order, because every other binary order reads only preserved data. In particular the profile and extension declarations are sorted in binary by the semantic keys $(\texttt{profile\_id}, \texttt{version})$ and $(\texttt{extension\_id}, \texttt{version})$, and the envelopes by canonical operation order. \end{requirement} \begin{rationale} Two bundles that differ only in physical layout are the \emph{same document}, and Requirement~\ref{req:textproj:canonical-text} obliges them to project to byte-identical text. Inheriting the binary order for these two sequences would let a chunk's file offset decide the order of the text --- so relocating a chunk, which changes no semantics, would change the projection. That is precisely the failure the requirement forbids. The de-duplication is the same point from the other side. A chunk is content-addressed: two entries with identical kind, schema version, and payload \emph{are} one chunk, and appear twice only because the writer stored the bytes twice. A blob with identical media type, declared maximum, and payload is one blob. Their projected forms are identical, so the binary form's distinction between them is a physical fact --- and a duplicated line would smuggle that physical fact into a text that claims to have erased it. Emitting one line is not a loss; it is the erasure working. Ordering by the projected form is total and deterministic, and it reads nothing but what is projected. For chunks it coincides with ordering by the derived \texttt{ChunkId}, since that id is a function of exactly the kind, schema, and payload the line carries --- which is a pleasing check that the rule is reading the right thing. \end{rationale} \section{Projecting Canonical Values} \label{sec:content:values} An operation payload embeds canonical values from the core specification's Chapter~5 --- an \texttt{Event}, a \texttt{Pitch}, a \texttt{Region}, a \texttt{TimeSignature}. This document does \emph{not} restate their shapes. It states one rule for turning any of them into text, and the shapes stay where they are ratified. \begin{requirement} \label{req:textproj:operation-vocabulary} The core specification's Chapter~6 operation vocabulary --- the envelope, its stamp and causal context, the payload, the operation kinds, and their sub-vocabularies --- \MUST{} be projected by the productions of Chapter~\ref{ch:grammar}, not by Requirement~\ref{req:textproj:value-projection}. The value-projection rule governs exactly the \texttt{value} positions those productions contain, and no other positions. An operation kind's payload record \MUST{} be inlined into its production. The record exists so each variant can name a type; it is not a modelling distinction, and the binary form agrees: \texttt{OperationKind}'s encoding writes the kind tag and then delegates to the record, adding no bytes for the wrapper. It adds no text here either, for the same reason clause~2 of Requirement~\ref{req:textproj:value-projection} makes a newtype transparent. \end{requirement} \begin{requirement} \label{req:textproj:schema-directed} The projection is \textbf{schema-directed}. At every position a reader knows the type it expects, from the ratified schema and from the position itself, exactly as the binary decoder does. The grammar of Chapter~\ref{ch:grammar} describes the \emph{shape} of the text; it is not a standalone unambiguous language, and a reader \MUSTNOT{} attempt to recover a value's type from its shape. Three consequences follow, and a reader resolves each by the type it expects: \begin{itemize} \item \texttt{()} is the empty sequence, and it is the absent option. \item A bare symbol is a fieldless variant, and it is a struct with no fields. \item A parenthesised list whose first element is a symbol is a struct, and it is a sequence whose first element is a symbol. \end{itemize} \end{requirement} \begin{rationale} This states what the ratified rules already require rather than adding a new constraint. A struct is \texttt{( \ldots)} and a sequence is a parenthesised list of its elements; a sequence whose first element is a fieldless variant is therefore shape-identical to a struct. The collision is \emph{irreducible} without new syntax, and it is reachable from the first pitched note a score contains: a \texttt{PitchedEvent} carries \texttt{articulations} and \texttt{ornaments}, sequences of a zero-field \texttt{ArticulationMark} and \texttt{OrnamentMark}, alongside an optional \texttt{DynamicMark} --- so one \texttt{insert-event} line holds an empty sequence and an absent option, both spelled \texttt{()}, and a two-element sequence spelled like a two-field struct. Making the text standalone-unambiguous would buy nothing, because Requirement~\ref{req:textproj:strict-parse} already obliges a parser to reject ``a duplicate in a set-typed field'' --- which it cannot do without knowing that the field is set-typed. The parser consults the schema either way. Spending syntax to remove a shape ambiguity, while leaving the schema dependency it was meant to remove, would make every line longer and no tool simpler. The binary form is schema-directed for the same reason and pays the same price: its bytes do not say what they are either. \end{rationale} \begin{requirement} \label{req:textproj:value-projection} A canonical value is projected thus: \begin{enumerate} \item A \textbf{struct} becomes \texttt{( \ldots)}, where the type name is the core specification's name in lower-case hyphenated form and the fields appear \emph{positionally}, in the order that specification's ratified listing declares them. Field names are not written. A struct with \emph{no} fields is the bare symbol \texttt{}, as a fieldless variant is; it encodes to no bytes in the binary form and carries no value here. \item A \textbf{newtype} --- a struct of exactly one unnamed field --- is projected as that field alone, with no wrapper. This mirrors the binary form, in which a newtype delegates to its field and adds no bytes. \item A \textbf{tagged union} becomes \texttt{( \ldots)}. A variant with no fields is the bare symbol \texttt{}. \item An \textbf{option} is \texttt{()} when absent and \texttt{(some )} when present. \item A \textbf{sequence}, \textbf{set}, or \textbf{map} is a parenthesised list of its elements, \emph{in the order the binary form writes them}, except where Requirement~\ref{req:textproj:derived-ordering} applies; a map entry is \texttt{( )}. An empty one is \texttt{()}, which Requirement~\ref{req:textproj:schema-directed} distinguishes from an absent option by the expected type. The projection invents no ordering of its own, and a set that the binary form writes strictly increasing is written strictly increasing here (Requirement~\ref{req:textproj:strict-parse}). A \textbf{byte string} is not a sequence. Where the binary form writes a length-prefixed run of bytes rather than a counted sequence of elements --- an opaque extension payload, a \texttt{SoundConfiguration} --- the projection writes a byte string, not a list of integers. \item \textbf{Leaves.} An identifier or hash is a byte string. An integer is an integer. A boolean is \texttt{true} or \texttt{false}. Canonical text is a quoted string. A rational is \texttt{(ratio )}, in lowest terms with a positive denominator and the sign on the numerator; zero is \texttt{(ratio 0 1)}. A \texttt{CanonicalF64} is a byte string of its eight canonical little-endian IEEE~754 bytes. \end{enumerate} \end{requirement} \begin{rationale} \textbf{One rule, not forty productions.} Spelling out a production per value type would restate the entire Chapter~5 data model in a second normative document, and two normative listings of one struct is precisely the drift this project has already been bitten by. A rule cannot drift from the listing it reads. \textbf{Positional fields.} The declaration order is already normative --- the binary form depends on it --- so field names would be redundant, would double the length of every line, and would create a second thing to keep in step with a rename. The constructor name carries the context a reader needs. \textbf{Newtypes are transparent} for the same reason they are transparent in the binary form: a \texttt{MusicalPosition} \emph{is} a rational, and wrapping it would put a distinction in the text that the document does not make. \textbf{Floats are bytes, never decimal.} A decimal rendering of an \texttt{f64} is not canonically unique --- shortest-round-trip and seventeen-significant-digit forms both round-trip, and $-0.0$ has two spellings --- so a decimal float would break Requirement~\ref{req:textproj:canonical-text} at the first tempo mark. The core specification already forbids computed floats in canonical state and stores the eight bytes; the projection carries those eight bytes. This costs readability in exactly one place, and buys canonicality everywhere. \end{rationale} \section{Reduced State} \begin{requirement} \label{req:textproj:reduced-state-derived} The projection \MUST{} preserve canonical reduced state \emph{by preserving the operations that determine it}. It \MUSTNOT{} carry a second, literal copy of the reduced state. \end{requirement} \begin{rationale} Reduced state is a deterministic function of the operation set and the canonical base (core specification Chapter~6, \sectionsc{Design Principles}). A text that carried both would have two sources of truth for one fact, and nothing could stop them disagreeing --- a projection with an internally contradictory document is worse than no projection. This reading is what the core specification's own round-trip clause already implies: text that ``parses to identical canonical document semantics'' must re-serialize to bundles with identical semantics, and reduced state is a function of semantics. Read the core specification's ``all canonical reduced state'' as \emph{determines}, not \emph{contains}. Ratified at 0.1.0. \end{rationale} \section{The Canonical Base Snapshot} \begin{requirement} \label{req:textproj:base-snapshot-inline} If the manifest declares a canonical base, the projection \MUST{} carry a \texttt{(canonical-base ...)} line bearing: \begin{itemize} \item the \texttt{SnapshotId}, verbatim --- the one identity that is carried rather than derived (Requirement~\ref{req:textproj:derive-or-carry}); \item the causal frontier it materializes, as opaque bytes; \item its reduction-algorithm version; \item the profile under which it was produced; \item its root chunk's \texttt{schema\_version}; and \item \emph{the root chunk's uncompressed payload, inline}, as a single byte string. That payload is the canonical byte form of the reduced state the snapshot materializes. \end{itemize} The root chunk's kind is \texttt{Snapshot} by role and is not written. The \texttt{SnapshotRef}'s \texttt{root} and \texttt{hash} are \textbf{re-derived}: the chunk's content hash is $\textrm{hash}(\texttt{Snapshot}, \textit{schema}, \textit{payload})$ under the Binary Format companion's chunk-hash preimage, and it is both the root's \texttt{ChunkId} and the \texttt{SnapshotRef}'s \texttt{hash}. A parser \MUST{} perform that derivation rather than read it from the text. \end{requirement} \begin{rationale} A canonical base exists precisely so that the operations before its frontier need not be retained. Where they have been pruned, the snapshot is \emph{not} derivable from anything else in the document, and a projection that carried only a reference would be \textbf{lossy} --- the text would no longer determine the document, which is the one thing it is for. The core specification permits the payload to be ``encoded compactly or referenced externally''; inline and compact is the choice that keeps the text self-contained. The snapshot diffs as one opaque atom. That is acceptable: merges happen among operations, and a base snapshot changes only when the document is compacted, at which point the whole line changes anyway. A later revision \MAY{} project the snapshot structurally; doing so does not break the round trip, because the document it denotes is unchanged. Ratified at 0.1.0. Schema major~0 has no snapshot producer --- pruning and canonical-base creation are deferred (Binary Format, \sectionsc{SnapshotId}) --- so this requirement binds whoever writes the first one, and cannot be exercised before then. \end{rationale} % =========================================================================== \chapter{Requirements} \label{ch:requirements} \section{Canonicality} \begin{requirement} \label{req:textproj:canonical-text} Two bundles whose canonical document semantics are identical \MUST{} project to \textbf{byte-identical} text. A projector \MUSTNOT{} have any freedom the document does not determine: no optional whitespace, no alternative spelling of an atom, no ordering choice. \end{requirement} \section{Round Trip} \begin{requirement} \label{req:textproj:roundtrip} Parsing a projection and re-serializing it to binary \MUST{} yield a bundle whose canonical document semantics are identical to the original's. The bundle's physical layout, chunking, and compression \MAY{} differ. \textbf{Documents carrying a canonical base are outside this requirement's domain, and are refused rather than round-tripped.} The exclusion is stated here, in the requirement's own terms, because the equations below quantify universally and a reader checking them against an implementation would otherwise find a conforming implementation failing them. The companion \MUST{} refuse a base-bearing document at every one of its three boundaries: projecting a base-bearing bundle to text, parsing text that declares a canonical base, and serializing a directly constructed document that carries one. The reason is that this format cannot carry the guarantee the base needs. A canonical base is only meaningful when it has been validated against a reduction authority, and that validation is attested by the \emph{container} epoch (Core Specification, \sectionsc{The Container Epoch}), which a text document has no way to hold: every field a text format defines can be typed by hand, so any provenance marker it carried would reduce to trusting its author. Serializing a base-bearing text document into a freshly created container would therefore mint a container asserting a validation that never occurred. A canonical base enters a container only through an explicit rebuild or repack flow, never through text import. This is a real loss of capability and is stated as one: a base-bearing bundle does not round-trip through text. The quantifiers below range over documents that carry no canonical base. Equivalently, and more usefully to an implementer: for every bundle $B$, \[ \textrm{semantics}(\textrm{parse}(\textrm{project}(B))) = \textrm{semantics}(B), \] and for every valid projection $T$, \[ \textrm{project}(\textrm{serialize}(\textrm{parse}(T))) = T . \] The second equation is the text's own injectivity: it is \emph{stronger} than the first, and it is the one a conformance test can check with byte equality. \end{requirement} \section{Strict Parsing} \begin{requirement} \label{req:textproj:strict-parse} A parser \MUST{} reject any text that is not the canonical projection of the document it denotes. It \MUSTNOT{} normalize: not whitespace, not letter case in a byte string, not an escape sequence, not an out-of-order sequence, not a duplicate in a set-typed field. Accepting non-canonical text and normalizing it \emph{is} accepting it, and does not satisfy this requirement. Rejecting a duplicate in a set-typed field requires knowing that the field is set-typed. A parser therefore reads the schema (Requirement~\ref{req:textproj:schema-directed}); shape alone never suffices. \end{requirement} \begin{rationale} This is the same discipline the binary decoders carry, and it exists for the same reason: a lenient parser makes two texts denote one document, and the projection's contract is that a text \emph{determines} its document. The Binary Format companion learned this concretely --- a whole-value re-encode guard catches the fields a decoder normalizes and is blind to order-preserving sequences, and a guard on an outer value can \emph{mask} a lenient inner codec rather than fix it (Binary Format, \sectionsc{The Decode Vector Corpus}). A text parser inherits both hazards, and the cheapest total defence is the same one: re-project the parsed document and compare, \emph{and} check per-site the orders that re-projection would restore. \end{rationale} \section{Conformance} \begin{requirement} \label{req:textproj:conformance} An implementation claiming Text Projection conformance \MUST{} implement both directions. A projector alone does not conform: the round trip (Requirement~\ref{req:textproj:roundtrip}) is the requirement, and half of it is not checkable. \end{requirement} % =========================================================================== \chapter{Grammar} \label{ch:grammar} \begin{lstlisting} ; Notation. X* is zero or more X, X? is zero or one. Within a line, adjacent ; elements of a repetition are separated by exactly one U+0020, and an empty ; repetition contributes nothing -- so "(" value* ")" spells () when empty. The ; line productions below carry their own trailing LF, and are simply ; concatenated. LF is U+000A. Terminals are quoted; U+XXXX names a codepoint. projection ::= header document lineage? profile* extension* canonical-base? blob* envelope* header ::= "(text-projection " version ")" LF version ::= "(" integer " " integer " " integer ")" document ::= "(document " bytes " " schema ")" LF ; id, then the carried manifest SchemaVersion (G-minor) lineage ::= "(lineage " bytes ")" LF profile ::= "(profile " profile-id " " version " " constraints ")" LF profile-id ::= "full" | "read-only" | "lite" | "(custom " bytes ")" constraints ::= "(constraints " integer " " retention ")" retention ::= "(retention " integer " " option " " bool ")" extension ::= "(extension " bytes " " version " " bool " (" chunk* ") " bytes " " bytes ")" LF ; id, version, required, chunks, affected-kinds, barriers ; (the ratified declaration order) chunk ::= "(chunk " chunk-kind " " schema " " bytes ")" chunk-kind ::= "operation-envelope-block" | "operation-index" | "snapshot" | "blob" | "extension-data" | "text-projection" | "layout-cache" | "integrity-index" | "manifest" schema ::= "(schema " integer " " integer ")" canonical-base ::= "(canonical-base " bytes " " bytes " " integer " " profile-id " " schema " " bytes ")" LF ; snapshot-id, frontier, reduction version, profile, ; root schema, root payload blob ::= "(blob " string " " option " " bytes ")" LF ; media type, declared max uncompressed length, payload envelope ::= "(envelope " bytes " " bytes " " stamp " " causal " " option " " payload ")" LF ; id, author, stamp, causal context, transaction, payload stamp ::= "(stamp " integer " " integer " " bytes ")" causal ::= "(causal (" replica-seen* ") (" bytes* "))" replica-seen ::= "(" bytes " " integer ")" payload ::= "(primitive " kind ")" | "(resolve-conflict " bytes " " action ")" | "(undo " bytes " " policy ")" | "(resolve-equivocation " bytes " " bytes ")" action ::= "accept-loser" | "keep-winner" | "dismiss" | "(override " bytes ")" | "(reanchor " bytes ")" | "(registered " bytes ")" policy ::= "strict-inverse" | "best-effort" | "cascade" ; --- Operation kinds inline their payload records as required by ; --- req:textproj:operation-vocabulary. Fields are the Operation Catalog's ; --- payload schema, positionally, in declaration order. Embedded Chapter-5 ; --- values occur only at positions and follow ; --- req:textproj:value-projection. kind ::= "(insert-event " bytes " " value ")" | "(delete-event " bytes " " tuplet-comp ")" | "(respell-pitch " bytes " " value ")" | "(create-cross-cutting " cross-cutting ")" | "(change-region-time-model " bytes " " value " (" bytes* ") " remapping ")" | "(set-user-system-break " bytes " " value " " bool ")" | "(declare-transaction " bytes " " string " " option ")" | "(registered " bytes " " bytes ")" | "(modify-event " value ")" | "(transpose (" bytes* ") " integer ")" | "(insert-identified-pitch " bytes " " value ")" | "(delete-identified-pitch " bytes ")" | "(modify-identified-pitch " bytes " " value ")" | "(delete-cross-cutting " bytes ")" | "(modify-cross-cutting " cross-cutting ")" | "(create-region " value ")" | "(delete-region " bytes ")" | "(create-staff-instance " bytes " " value ")" | "(delete-staff-instance " bytes ")" | "(create-voice " bytes " " value ")" | "(delete-voice " bytes ")" | "(set-metadata " value ")" | "(set-metric-grid " bytes " " option ")" | "(set-user-page-break " bytes " " value " " bool ")" | "(create-staff " value ")" | "(set-time-signature " bytes " " value " " option ")" | "(set-tempo-segment " option " " value " " option ")" | "(set-staff-layout " bytes " " option " " option " " bool ")" | "(create-repeat-structure " value ")" | "(delete-repeat-structure " bytes ")" | "(transpose-interval (" bytes* ") " value ")" | "(create-instrument " value ")" | "(set-canvas-layout-defaults " value ")" | "(set-spelling-precedence " value ")" | "(set-tuning-context " value ")" | "(create-staff-group " value ")" | "(create-part-definition " value ")" | "(create-analysis-layer " value ")" | "(create-view " value ")" | "(create-measure " value ")" tuplet-comp ::= "not-in-tuplet" | "(replace-with-rest " value ")" | "(rewrite-tuplets (" bytes* "))" | "(cascade-delete-tuplets (" bytes* "))" cross-cutting ::= "(tie " value ")" | "(slur " value ")" | "(beam " value ")" | "(spanner " value ")" remapping ::= "preserve-time" | "(reassign (" reassign-entry* "))" reassign-entry ::= "(" bytes " " ratio ")" ; event id, musical position ; --- Values and leaves. ; A struct, a sequence, a set, a map and an option all have this one shape. The ; reader tells them apart by the type it expects, never by the shape: ; req:textproj:schema-directed. Meaning is req:textproj:value-projection. value ::= "(" value* ")" | symbol | bytes | integer | bool | string | ratio | option option ::= "()" | "(some " value ")" ratio ::= "(ratio " integer " " integer ")" bytes ::= "#x" hexdigit* ; even count, lowercase integer ::= "-"? digit+ ; no leading zeros, no "-0" bool ::= "true" | "false" symbol ::= [a-z] [a-z0-9-]* string ::= '"' schar* '"' digit ::= [0-9] hexdigit ::= [0-9a-f] schar ::= unescaped | escape unescaped ::= escape ::= U+005C U+0022 ; the two characters \" | U+005C U+005C ; the two characters \\ | U+005C "n" ; the two characters \n | U+005C "t" ; the two characters \t \end{lstlisting} Observe what does \emph{not} appear: no offset, no compressed length, no compression algorithm, no chunk id, no content hash, no blob id. Every one is either physical or derivable, and Requirement~\ref{req:textproj:derive-or-carry} forbids both. The lone opaque identity the grammar carries is the \texttt{SnapshotId}. \textbf{Every production is expanded.} Version 0.1.0 left \texttt{kind}, \texttt{action}, \texttt{policy}, \texttt{constraints}, and \texttt{barrier} derived-but-unwritten, and said so; that gap is closed. Barriers and affected object kinds are byte strings, because the bundle holds them opaquely and the projection interprets nothing the bundle does not. At exactly the \texttt{value} positions delimited by Requirement~\ref{req:textproj:operation-vocabulary}, the one thing still \emph{read} rather than restated is a Chapter-5 value's field list. That is deliberate: Requirement~\ref{req:textproj:value-projection} is a rule applied to the core specification's ratified listings, not a copy of them. A rule cannot drift from what it reads. An implementation that projects a value's fields in an order other than the declaration order disagrees with the core specification, not with this document. Note the operation-kind names are the \emph{Operation Catalog's} section names (\texttt{create-region}, \texttt{create-staff}), not the \texttt{OperationKindTag} names (\texttt{InsertRegion}, \texttt{InsertStaff}). The tag space renamed three pairs for reasons of its own (Binary Format, \sectionsc{\texttt{OperationKindTag}}); the projection follows the semantics, not the tag. % =========================================================================== \chapter{A Worked Example} \label{ch:example} \emph{Non-normative.} The byte strings below are illustrative but \emph{well-formed}: every one is a grammar-valid \texttt{bytes} terminal, because an example that cannot be parsed teaches the wrong lesson. The conformance vectors that pin \emph{real} bytes are a deliverable of the implementation. A document of one operation --- a transposition of two pitches up a perfect fifth --- projects to four lines: \begin{lstlisting} (text-projection (0 14 0)) (document #x05050505050505050505050505050505 (schema 0 1)) (profile full (0 1 0) (constraints 67108864 (retention 1 () true))) (envelope #x00000000000000070000000000000001 #x00000000000000000000000011223344 (stamp 42 7 #x00000000000000070000000000000001) (causal ((#x0000000000000001 3)) (#x00000000000000020000000000000009)) (some #x00000000000000070000000000000005) (primitive (transpose-interval (#x00000000000000070000000000000001 #x00000000000000070000000000000002) (transposition-interval 4 7)))) \end{lstlisting} The envelope's targets are a \emph{set}: strictly increasing, no duplicates (Operation Catalog, \texttt{req:opcat:transpose-interval-targets}; Binary Format $\mathrm{seq}^{\Uparrow}$). A parser \MUST{} reject a duplicate rather than absorb it, exactly as the binary decoder does. Earlier revisions of this companion showed the same document over a compacted base, as five lines. The fifth line is retained here as the spelling a parser now \emph{refuses}: \begin{lstlisting} (canonical-base #x1f8b0000000000000000000000000000 #x 1 full (schema 0 1) #x0000) \end{lstlisting} The grammar still defines the section --- it is what a refusal is defined against, and what a future rebuild or repack flow will emit --- but no document containing it parses, projects, or serializes (\texttt{req:textproj:roundtrip}). % =========================================================================== \chapter{Revision History} \label{ch:history} \begin{longtable}{p{2cm} p{2.5cm} p{9cm}} \toprule \textbf{Date} & \textbf{Section} & \textbf{Change} \\ \midrule \endhead \today & All & 0.1.0 --- Initial companion. Supplies the canonical s-expression form the core specification's Chapter~8 \sectionsc{Text Projection} declares normative and leaves unwritten, and which the Binary Format companion excludes as ``the \emph{Text Projection} companion's''. Ratified: reduced state is preserved by \emph{determining} it, never by a second literal copy (\texttt{req:textproj:reduced-state-derived}); a canonical base snapshot is inlined as one opaque byte string, because a pruned document's base is derivable from nothing and a reference-only projection would be lossy (\texttt{req:textproj:base-snapshot-inline}); lowercase hex is the single byte-string encoding (\texttt{req:textproj:hex}); one envelope per line, so a line-based merge conflict is exactly an envelope conflict (\texttt{req:textproj:envelope-per-line}). Parsing is strict-canonical (\texttt{req:textproj:strict-parse}) --- normalizing non-canonical text is accepting it --- and conformance requires \emph{both} directions (\texttt{req:textproj:conformance}). No implementation yet; this document is the design gate. \\ \today & Chapters 3, 5 & 0.2.0 --- Canonical-manifest coverage. 0.1.0 was \emph{lossy for documents that are valid today}, and its claim to preserve the manifest's canonical roots was false in three ways: a canonical blob (an embedded image, font, or recording referenced by a canonical operation) had no representation at all; an \texttt{ExtensionDeclaration} lost its semantic version and its affected object kinds, and left its preserved chunk roots undefined; and \texttt{ProfileId::Custom} was unrepresentable, a symbol being required where a sixteen-byte registry id is carried. The three share one cause, now stated as \texttt{req:textproj:derive-or-carry}: a \texttt{ChunkRef} and a \texttt{BlobRef} are \emph{physical} references --- offset, compressed length, compression --- which the projection may not preserve, and they also carry \emph{derivable} identities, which it may not duplicate. Carry the content and the semantic attributes; re-derive the rest. The sole non-derivable identity in schema major~0 is \texttt{SnapshotId}, which the Binary Format companion pins as opaque and forbids readers to derive. Consequently: \texttt{req:textproj:canonical-blobs} (a canonical blob is projected; a non-canonical one is not), \texttt{req:textproj:profile-id} (\texttt{(custom \#x...)}), \texttt{req:textproj:extension-declaration} (every field, chunks as kind + schema + payload), and \texttt{req:textproj:base-snapshot-inline} extended to say what the inlined payload \emph{is} and how the root chunk id and the snapshot hash are re-derived rather than read. The core specification's own list of what the projection preserves omitted canonical blobs; it is corrected there too. The 0.1.0 ratifications --- reduced state derived, base inlined, hex, one envelope per line, strict parsing --- stand unchanged. \\ \today & Chapters 3, 5 & 0.3.0 --- Every production expanded. 0.1.0 left \texttt{kind}, \texttt{action}, \texttt{policy}, \texttt{constraints} and \texttt{barrier} derived-but-unwritten and admitted it; all are now written. Barriers and affected object kinds are byte strings, because the bundle holds them opaquely and the projection interprets nothing the bundle does not. Operation payloads embed Chapter-5 canonical values, and those are projected by \emph{one rule} rather than by forty productions (\texttt{req:textproj:value-projection}): a struct is \texttt{( \ldots)} with fields positional in the ratified declaration order; a newtype is transparent, as it is in the binary form; a tagged union is \texttt{( \ldots)}; an option is \texttt{()} or \texttt{(some v)}; a sequence keeps the binary form's order. Restating the Chapter-5 model here would have put two normative listings on one struct, which is the drift P13-I1 was opened to close --- and a rule cannot drift from what it reads. Two leaf decisions follow from canonicality rather than taste. A rational is \texttt{(ratio n d)} in lowest terms with the sign on the numerator. A \texttt{CanonicalF64} is the byte string of its eight canonical IEEE~754 bytes, never a decimal: decimal float text is not canonically unique, so a decimal tempo would break \texttt{req:textproj:canonical-text} at the first tempo mark. Operation-kind names are the Operation Catalog's section names, not the \texttt{OperationKindTag} names, which renamed three pairs for reasons of the tag space. Still no implementation. \\ \today & Chapters 3, 5 & 0.4.0 --- Two normative corrections found in review. \emph{The grammar contradicted its own escape requirement.} \texttt{req:textproj:string-escapes} obliges a writer to escape the backslash and a parser to reject a bare one, while \texttt{unescaped} admitted it. The escape productions are now spelled out as two-character sequences and \texttt{unescaped} excludes U+0022, U+005C, U+000A, and U+0009 by codepoint. Both characters of an escape are written as codepoints where they are the delimiter or the introducer: a quoted terminal for the backslash reads as \emph{two} backslashes and would make every escape three characters long. For the same reason the mono font no longer applies TeX ligatures, which rendered U+0022 as a right curly quote and \texttt{-{}-} as an en dash --- a document that specifies a text syntax must not misprint it. \emph{``Keep the binary order'' does not work for every sequence.} It is available only where the binary order reads preserved data, and two sequences fail that test: \texttt{blob\_roots} is sorted by the full \texttt{BlobRef} encoding, which contains the offset, the compressed length, and the compression; and an extension's preserved chunk roots are sorted by \texttt{ChunkRef}'s order, whose key is kind, then content hash, then \emph{offset}. Under the blanket rule, relocating a chunk --- which changes no semantics --- would have changed the text, and two entries indistinguishable after erasure would have produced duplicate lines. \texttt{req:textproj:derived-ordering} orders and de-duplicates those two sequences by their \emph{projected form}. Every other sequence keeps the binary order: the profile and extension declarations sort on semantic \texttt{(id, version)} keys, and the envelopes on canonical operation order. Also: a committed grammar-completeness test now checks that no nonterminal is undefined or unreachable, that the escape rule excludes the four codepoints and admits exactly the four two-character sequences, that the mono font substitutes no glyphs, and that the operation-kind and chunk-kind productions are exactly the tag vocabularies --- derived from \texttt{OperationKindTag} and \texttt{ChunkKind}, not transcribed. Every locator in it finds its production by name; the checks it replaced were anchored to a column, and a reflow would have silently switched them off. The 0.3.0 claim of a ``machine-checked'' grammar was true of one run and of nothing durable. Still no implementation. \\ \today & Chapters 3, 5 & 0.5.0 --- Found by starting the implementation: the grammar could not derive an ordinary pitched note. \emph{The projection is schema-directed} (\texttt{req:textproj:schema-directed}), and now says so. \texttt{value} had no alternative for a sequence at all, though clause~5 required one; and once added, a sequence whose first element is a fieldless variant is shape-identical to a struct, while \texttt{()} is both the empty sequence and the absent option. Both collisions are reachable from the first pitched note in any score --- \texttt{PitchedEvent} carries \texttt{articulations} and \texttt{ornaments} over zero-field marks, beside an optional \texttt{DynamicMark}. The collision is irreducible without new syntax, and new syntax would buy nothing: \texttt{req:textproj:strict-parse} already obliges a parser to reject a duplicate in a set-typed field, which it cannot do without the schema. So \texttt{value} now collapses to \texttt{"(" value* ")"} or a leaf --- all shape can honestly say --- and the requirement assigns meaning by expected type. \emph{Three consequences, stated.} A struct with no fields is the bare symbol, as a fieldless variant is. A byte string is not a sequence: where the binary form writes a length-prefixed run of bytes, so does the projection, never a list of integers. And the grammar's repetitions now carry a notation rule --- adjacent elements are separated by exactly one space --- without which \texttt{"(transpose (" bytes* ") "} spelled two targets as one run of hex. Still no implementation; this is what the first day of writing it found. \\ \today & Chapters 3, 5, 6 & 0.6.0 --- The operation vocabulary is grammar-directed, and the boundary is now normative (\texttt{req:textproj:operation-vocabulary}). The envelope, stamp, causal context, payload, operation kinds, and their sub-vocabularies follow the grammar's productions; \texttt{req:textproj:value-projection} applies exactly at the \texttt{value} positions those productions name. Operation-kind payload records are inlined because they name variant types but add neither a binary wrapper nor a modelling distinction. \texttt{transpose-interval} now takes its target byte strings followed by one \texttt{value}. Its \texttt{TranspositionInterval} therefore has the same \texttt{(transposition-interval ...)} spelling it has at every other value position, and the grammar has no special \texttt{interval} production. The committed grammar gate locks the requirement, its reliance-point citations, and that absence. \\ \today & Chapters 3, 5 & 0.7.0 --- Canonical blob rejection and single-version header gating. At this version no canonical operation or canonical reduced state can reference a \texttt{BlobId}; every \texttt{(blob ...)} line is therefore unreferenced and non-canonical, and a parser must reject it (\texttt{req:textproj:reject-unreferenced-blobs}). Accepting it would stage content the next projection silently drops, lose data, and falsify the text-to-binary-to-text round trip. Forward compatibility belongs to header-version gating, not lenient blob acceptance. The parser accepts exactly the implemented companion header, \texttt{(0 7 0)}, and rejects every other version (\texttt{req:textproj:header-version}). Multi-version acceptance and text migrate-on-read remain deferred, in the same posture as op-payload migrate-on-read: a future real consumer must bring an explicit, version-keyed migration path rather than teaching the current parser to speculate. The worked example now uses the implemented header and contains only grammar-valid, canonical lines. \\ \today & Chapter 5 & 0.8.0 --- The genesis operation vocabulary reaches the grammar. The \texttt{kind} production gains \texttt{"(create-instrument " value ")"}, the first operation kind appended since the header was gated to a single version at 0.7.0 (\texttt{req:textproj:operation-vocabulary}). The bump is forced rather than cosmetic. A parser \MUST{} accept exactly the version of the companion it implements (\texttt{req:textproj:header-version}), so extending the grammar while holding the version would leave two mutually incompatible grammars both claiming \texttt{(0 7 0)} --- an older parser meeting the new production would fail with no version signal to explain why, which is the precise failure the single-version gate exists to prevent. Cached projections at \texttt{(0 7 0)} do not migrate and are not expected to. A \texttt{TextProjection} chunk is a non-canonical accelerator a writer may discard, so a stale projection is regenerated from the canonical document rather than converted. Multi-version acceptance and text migrate-on-read remain deferred, unchanged. \\ \today & Chapter 5 & 0.9.0 --- The genesis settings setters reach the grammar. The \texttt{kind} production gains \texttt{"(set-canvas-layout-defaults " value ")"} and \texttt{"(set-spelling-precedence " value ")"} (\texttt{req:textproj:operation-vocabulary}), the second appended pair since the header was gated to a single version at 0.7.0. Same forcing reason as the 0.8.0 bump above: holding the version while extending the grammar would leave two mutually incompatible grammars both claiming \texttt{(0 8 0)}. Cached projections at \texttt{(0 8 0)} do not migrate; a stale \texttt{TextProjection} chunk is regenerated, not converted. \\ \today & Chapters 8, 5 & 0.10.0 --- The carried manifest schema version reaches the \texttt{document} line (the G-minor rung, \texttt{spec/PLAN\_GMINOR\_SCHEMA\_MINOR.md}). The \texttt{document} production gains a second field: the manifest's aggregate \texttt{SchemaVersion}, carried verbatim and never derived (\texttt{req:textproj:manifest-schema-carried}). This bump is caused by the manifest attribute alone, \emph{not} by operation-block schema-minor stamping: an operation block's physical schema is discarded during projection exactly as before, and a reader must not infer that op-block minors reach the text surface from this entry. The reasoning is the same forcing argument as every prior grammar-extending bump: holding the version while adding a field would leave two mutually incompatible grammars both claiming \texttt{(0 9 0)}. Cached projections at \texttt{(0 9 0)} do not migrate; a stale \texttt{TextProjection} chunk is regenerated, not converted. \\ \today & Chapter 5 & 0.11.0 --- The genesis tuning-context setter reaches the grammar (genesis tranche G2b, \texttt{spec/CONTRACT\_GENESIS\_G2B\_TUNING.md}). The \texttt{kind} production gains \texttt{"(set-tuning-context " value ")"} (\texttt{req:textproj:operation-vocabulary}), the third appended kind since the header was gated to a single version at 0.7.0. Same forcing reason as every prior grammar-extending bump: holding the version while extending the grammar would leave two mutually incompatible grammars both claiming \texttt{(0 10 0)}. Cached projections at \texttt{(0 10 0)} do not migrate; a stale \texttt{TextProjection} chunk is regenerated, not converted. \\ \today & Chapter 5 & 0.12.0 --- The four remaining genesis root-level entity mints reach the grammar (genesis tranche G3a, \texttt{spec/CONTRACT\_GENESIS\_G3A\_ENTITIES.md}). The \texttt{kind} production gains \texttt{"(create-staff-group " value ")"}, \texttt{"(create-part-definition " value ")"}, \texttt{"(create-analysis-layer " value ")"}, and \texttt{"(create-view " value ")"} (\texttt{req:textproj:operation-vocabulary}), the fourth appended kind event since the header was gated to a single version at 0.7.0. Same forcing reason as every prior grammar-extending bump: holding the version while extending the grammar would leave two mutually incompatible grammars both claiming \texttt{(0 11 0)}. Cached projections at \texttt{(0 11 0)} do not migrate; a stale \texttt{TextProjection} chunk is regenerated, not converted. \\ \today & Chapter 5 & 0.13.0 --- The genesis ladder closes: \texttt{create-measure} reaches the grammar (genesis tranche G3b, \texttt{spec/CONTRACT\_GENESIS\_G3B\_MEASURE.md}). The \texttt{kind} production gains \texttt{"(create-measure " value ")"} (\texttt{req:textproj:operation-vocabulary}), the fifth appended kind event since the header was gated to a single version at 0.7.0. Same forcing reason as every prior grammar-extending bump: holding the version while extending the grammar would leave two mutually incompatible grammars both claiming \texttt{(0 12 0)}. Cached projections at \texttt{(0 12 0)} do not migrate; a stale \texttt{TextProjection} chunk is regenerated, not converted. \\ \today & Chapters 4, 5 & 0.14.0 --- Canonical bases leave the companion (\texttt{spec/CONTRACT\_FORMAT\_EPOCH\_MAJOR1.md}, pin 3b). The container format major becomes an \emph{epoch} attesting that every base-bearing commit was validated against a reduction authority (Core Specification, \sectionsc{The Container Epoch}); a text document cannot hold that attestation, because every field this format defines can be typed by hand. Serializing a base-bearing document into a freshly created container would therefore mint a container asserting a validation that never happened. \texttt{req:textproj:roundtrip} accordingly excludes base-bearing documents in its own terms, and all three boundaries refuse them: projection, parsing, and serialization. The first two of those refusals are \emph{new}; serialization had none to retain. This is a real capability loss and is recorded as one --- a base-bearing bundle does not round-trip through text until a rebuild or repack flow exists. The forcing reason for the version bump is the usual one in reverse: refusing a document the companion previously serialized is a semantic change, and holding the version would leave two mutually incompatible readings of \texttt{(0 13 0)}. The \texttt{canonical-base} production is retained in the grammar --- a refusal is defined against it, and a repack flow will emit it. \\ \bottomrule \end{longtable} \end{document}