Architecture
The same library from underneath. Nothing here is needed to use it — this is for deciding whether it fits, reading a stack trace, or choosing an engine on something other than taste.
The parts, and what talks to what
Section titled “The parts, and what talks to what”The reactive layer is the store’s own code: it builds the handles, keeps the list of who is watching what, and calls them back. It is not a tier between the store and the handles, which is why the box is drawn inside.
The lit path is one write, and both halves of it happen before set returns.
The handle writes in; the subscription the layer registered when it built that
handle is what sets the handle’s own value — with the store’s stamp rather than
a fresh one, so two racing writes are ordered once and everyone sees the same
order. An edit that arrives from the file rather than from a caller comes back
along the very same edge.
Durable is not a component of its own — it is what any of the four handles
gets from .durable(). Both of its edges point both ways, and that is the whole
of what it does: it tells the debouncer not to wait its window out, and it blocks
until the commit answers. The answer comes back the same way, to the handle that
asked. Every other write on this diagram returns before the disk has been
touched.
The edge coming back off the disk is a text engine only. Those three watch their file because they are not its only writer, and what they read cannot go straight to a handle: a file arrives whole, and which paths in it changed is something that has to be worked out. That is the diff, and it is written on the edge rather than drawn as a box because that is what it is — a pass over two documents on the way in, not a part with a life of its own. What it works out is fired through the very same subscription list a local write goes through, which is why an application never learns where a change came from unless it asks. redb and SQLite have no such edge; their file is theirs.
Migrations come in through the store’s own API — run_migrations, which every
engine implements — but through a door of their own, and they are the one write
on this diagram that does not go through the debouncer. A step sees raw bytes
through the migration adapter, the whole pass runs in one transaction the engine
opens, and it lands when that transaction commits. There is nothing to debounce:
this happens while the store is opening, before a single handle exists, and the
next thing to run has to see it.
A path is a list of levels
Section titled “A path is a list of levels”A path is levels, and only levels. StorePath::segment("dark.mode") is one
level whose name happens to contain a dot. The dotted spelling is how a path
writes itself down, not what it is: joining escapes a separator or a backslash
inside a name, and reading it back undoes that. Everything that compares,
orders or hashes a path does so by its levels, so the two spellings of one name
can never be mistaken for each other.
A path keeps whichever form it arrived in and works the other out only if somebody asks. A declared path arrives from the macro with both halves already written into the binary and checked against each other at compile time; a path read back off a disk arrives in whichever form its engine holds, and a scan that never needs the other half never pays for it.
The flat engines address by an encoding rather than by the spelling: each level’s bytes, each level terminated. Because the terminator cannot appear inside a level, byte order over keys is level order and a subtree is exactly a byte prefix — so a prefix scan is one range over the table with nothing to filter afterwards. That is one decision, made once, and both flat engines get their scans from it.
Two families of engine
Section titled “Two families of engine”redb and SQLite are flat. One key-value space; a struct at a.b is one key
and a.b.x beside it is another. Values are msgpack on redb and JSON on SQLite.
Nothing has to be told which paths are structure and which are data, because
the key says so.
JSON, TOML and RON are documents. A document writes an object either way, so
it cannot tell “a struct at a.b” from “a level a holding a level b” — and
it has to, because a person opens this file. So the declarations decide: a
declared place, and every level on the way to one, is written as a nested tree
the way serde would have written it. Everything else — including a path inside
a declared value — goes in a plane of whole keys beside the tree, one level
named by the whole path. A map’s entries nest; a key nobody declared does not.
The three text engines share one traversal and one diff. What differs between them is only what a single node can hold, which is the format’s business and shows up as the limits each one has.
A write goes down and comes back up
Section titled “A write goes down and comes back up”field.set(x) does not update the field and then persist it. It hands the value
down, and the field learns about it on the way back.
Interceptors run first and may turn the write down. Then the value is encoded and compared against what is already stored: identical bytes stop here — no event, no disk, no subscriber woken. Otherwise it enters the buffer, and the store calls every matching subscriber synchronously, on the thread that wrote, before scheduling anything. The field’s own signal is set by that callback, with the store’s stamp rather than a fresh one, so two racing writes are ordered once and everyone sees the same order.
This is why set then get returns the new value, why an unreadable value
leaves the old one in place and reports itself instead, and why an edit made
outside the process reaches subscribers by exactly the same path a local write
does — differing only in where it says it came from.
Writes reach disk once per window. The first write after a flush opens it, the ones that follow join it without moving its end, and when it closes one thread flushes what accumulated: a burst becomes one flush, and writes that never pause still reach disk every window. A flush that fails is retried, and the retry budget bounds how long the store stays quiet about it, not how long it keeps trying — it keeps trying until it lands or the store is dropped. Outliving the budget reports once, hands back every path written since the last flush that landed, and asks what writers should be told from then on.
Closing writes down whatever was still held. A write that cannot wait asks for the disk directly — and what that buys differs by family: on a text engine one commit covers the whole store, so everything buffered becomes durable with it; on redb and SQLite it covers that write.
The disk, and the second writer
Section titled “The disk, and the second writer”A text store is two files — the data and what the library knows about it — each replaced through a temporary file in the same directory, flushed, then renamed. Each replacement is atomic on its own. The pair is not, and the order is chosen so that the disagreement a crash can leave is the one that can be detected: the bookkeeping claims more than the data holds, which the next open notices.
A text store also watches its file, because it is not the only writer. A touched file is not a changed one, so the first thing a look does is hash the bytes and compare them against what this store last wrote — no parse, no diff, and that is the common case. Bytes that really differ are parsed, and if this store is holding writes the file never got, they are laid back over what the file now holds rather than overwriting it. What the file brought with it is diffed and delivered as ordinary events.
What sits beside the data
Section titled “What sits beside the data”The version each prefix has reached, the shape each struct had when it last wrote, the log of steps applied, which namespaces have had their defaults written, and a record of how the bytes themselves were written.
That last one is the one that refuses. It names the codec, the key encoding and the document layout, and an open that meets a value it does not write stops and says which fact it stopped on. A name outside those namespaces is carried through untouched, so an older build writing a store a newer one made does not erase what it could not read.
The recorded shapes are compared by the places a declaration owns, never by the name of the struct or the type. Renaming a type is not a change to the store, and must not be read as one.
Migrations are one transaction per pass
Section titled “Migrations are one transaction per pass”A pass starts at a prefix and opens one transaction covering it and everything it reaches. Failure rolls that back and leaves the rest of the store alone, so one prefix’s bad step is something the report can name rather than something that stops the open.
Nothing declares an order. A step that reads across into another prefix brings that prefix up to date first, inside the same transaction — so the order is the reaching, and a cycle is caught while it is happening, before anything has been written that would need undoing.
Who owns a path
Section titled “Who owns a path”Three mechanisms, and they answer different questions. What a constructor in this process has already built refuses a second one and can say which instance took it. What the binary declares refuses a raw write, and it knows about structs nobody has built yet. What a previous run recorded on disk can only tell — a claim by code that is not running protects nobody, and is there to report drift rather than to forbid.
A fourth check sits outside all three, at read time: a map refusing an entry more than one level below itself. It is deliberately not part of the tables, because it is the only one that works against a writer no table knows about — a raw write, a migration, or a person with an editor.
What the shape costs
Section titled “What the shape costs”A path and a value is the whole data model: no queries, no indexes, no transaction you can open yourself.
A process holds as many stores as it likes — that is ordinary, and the per-instance bookkeeping exists for it. What the engines do not agree on is two stores over one file: redb refuses the second outright, SQLite holds the file, and the text engines let both in and reconcile at save time, which is what the standoff above is for. Whether the two are in one process or two makes no difference to any of them.
And a text engine’s file is meant to be read and edited by a person, which is why it costs a whole subsystem the flat engines do without.