Imported from zeroaltitude/theseus (
crates/theseus-store/AGENTS.md). Install upstream withnpx skills add zeroaltitude/theseus --skill theseus-store. Copyright stays with the author.
theseus-store
The keel (spec §6, Part II M1): an append-only WAL of checksummed, length-prefixed atomic frames, which is the truth, and a redb index rebuilt from it. Read by theseus-kernel, theseus-core, theseusd, theseus-sim, and theseus-exam, and by the reserved theseus-follow and theseus-index.
Key modules: wal.rs, index.rs, record.rs, store.rs (MANIFEST_FORMAT). Read by: kernel, core, theseusd, sim.
What's here
wal.rs: segment files of atomic frames, each with its synced mark, and torn-tail truncation.index.rs: the rebuildable redb projection, andmove_asidefor an index that is not a database.record.rs: records and their kinds, and the header's frozen schema field (FROZEN_SCHEMA).store.rs: theStorecontract the kernel writes through, andWalStore, which composes the WAL and the index, with its writer thread.MANIFEST.jsonnames the store's one format number (MANIFEST_FORMAT).blockingruns a wait for the disk without holding a runtime worker.pressure.rs(theseus-tood): a background pass waits between two chunks while the machine is busy (PSI'ssome avg10at or over the gate's 20 % CPU or 10 % IO, up toBOUND; nothing waits without PSI), andidle_this_threadputs a thread inSCHED_IDLE, one-way, so only one that answers no one and starts no thread others use (a thread inherits its creator's policy).wordsistheseusd check'sbackground:line.
Invariants
-
One writer (theseus-vni9). Every append goes through
WalStore'sstore-writerthread: the caller hands it the frame and waits (inblocking, so no runtime worker waits), and the writer writes every frame queued, syncs once for all of them, indexes them in one transaction, and answers each. The answer comes after the index, so a caller's lock spans its read to its frame indexed (K1). Never append from the writer itself, and never holdappendingwhile you append: the writer needs it shared. -
A segment's name is as durable as its frames (theseus-xprd). A roll syncs the segment it leaves; the
syncthat makes a new segment's first frame durable then syncs the log's directory before it returns, and a new log's first sync syncs the directory holding the log too. An open that appends to a last segment it found syncs the log's directory with its first frame as well, once (theseus-c67g): that segment's creator may have died before its own sync did, and ext4's ordered mode hiding it is no promise of POSIX's. It skips that sync when the open knows a position in the found segment was synced (theseus-3q29): the index's checkpoint, or a frame's mark, at or past the segment's first position, says a sync covering its first frame returned Ok, and such a sync syncs the pending directories before it returns.Recovery::vouchedsays which; an empty found segment is never vouched for. A store whose manifest names an older format (behind) syncs the log's directory with its first frame anyway, once, as the manifest moves: its marks may be a build's from before c67g, which synced no found segment's name. A file made durable needs its directory synced too. The store's open adds the holder of every directory it created, the store's own included (theseus-gf00), to those the first frame's sync makes durable. -
A sync that fails takes its frames with it (theseus-ljgm). Its batch is answered failed, so
Wal::synccuts the segment back to the end of the last frame a sync that returned Ok covered, syncs the cut, and rolls the writer back: the next frame takes the first cut position, andsynced(the mark) is left as it was, so no later sync's Ok (fsyncgate) claims a page the failed one may have dropped. Syncs run one at a time (Wal::durable), and a roll's sync is one of them. The log is broken instead (it refuses frames, saying why) when the cut or its sync fails, or when the open found frames past the last position known synced and no sync has covered them yet: their writer may have answered them. Followers meet the cut as a rewind (theseus-follow checks the frame behind its cursor after each read). An index write that fails after a good sync still fails its batch, and its frames, durable, come back at the next open: not yet fixed. -
The version rule: one format number (P5b; Part III F4a; theseus-ptx1, Tier 7's 7.9 as the owner amended it). Any step that adds a field to a stored record (nested ones included), or changes the frame or record encoding, bumps
MANIFEST_FORMAT, so an older binary refuses the newer store. It lands with the reader for the layout it replaces (serde defaults, or a reader such asExecution::from_stored) and adds a sample of that layout, as literal bytes its build wrote, to theseus-core'stests_layoutswhen the layout is on disk somewhere. A new record kind needs no table. The number itself is assigned when the step lands onmain, never in a lane: a lane that adds a field says so in its report. -
A record's header schema is frozen:
NewRecordhas none, the WAL writesFROZEN_SCHEMA(0), and nothing reads the field. Records from before one store format keep their kind's old number there. -
Old layouts are read in place. An open writes no manifest. The writer's first frame into a store an older build wrote moves its manifest to this build's format first, durably, once: two syncs, and on a daemon that frame is its kernel's startup frame, so the first start after an upgrade that bumps the format pays them before serving. A store only read keeps its format. A build older than the store refuses it, before anything is written ("install the newer theseusd"). There is no rolling back: keep the newer binary, or restore a copy taken before the upgrade.
-
The open reads only the WAL's tail, from the frame after the index's checkpoint. The rest is checked after serving by core's
store-verifythread, and a corrupt frame there is refused and loud. -
The open makes nothing durable (theseus-ptx1): the tables' creation is a non-durable commit, and a replay takes no checkpoint. The WAL is the truth, so a crash before the next checkpoint only replays the tail again; the replayed tail counts toward the next periodic checkpoint, which the writer takes after serving. A start then pays only redb's own sync at its open, and a start after a crash no repair when the run made no durable commit.
-
A list read skips a refused record (R4, theseus-15g):
read_manyandscanleave out a record whose read is refused (a corrupt frame), log it once, and count it inStoreStats::refused_records, which health shows. A read of that record alone (get,latest_by_key) is still refused.repair.rstakes a frame that does not check whole from a copy of the store, andtheseus_core::restore::repairswaps the repaired WAL in. -
An open cuts a bad frame in the last segment as a torn tail only past every position known synced (theseus-gt12, theseus-7nfj). Two facts say a position was synced, and the larger decides:
- the index's checkpoint, which claims only synced positions (it takes
appendingalone, and the writer indexes a batch only after its sync). The store's open passes it asWalConfig::synced_to, and a repair passes the checkpoint of the store it repairs (index::checkpoint_of, read-only); - the mark of any whole frame after the bad one: the writer stamps each frame with the last position whose
sync had returned Ok when it wrote the frame (
Wal::synced, advanced only by asyncthat returned Ok), so a later batch's frame proves the bad one synced. It needs no index: a full replay andtheseusd restorerefuse by it too.
A bad frame at or before either is refused, naming the evidence and
theseusd restore --repair, and nothing is cut. Past both, the bytes cannot tell rot from a batch torn before its sync (a later frame of the batch can reach the disk whole, and carries the batch before's mark), so the frame is cut, with all after it, andRecovery::cutsays whether a whole frame followed and how far the log was known synced;StoreStats::cut, health's store phase, and restore's report andstore.restoredrow carry it. Both walks (every segment, and the tail after the checkpoint) decide alike. What this still cuts: rot in the log's last batch, until a checkpoint reaches it, and rot in an unmarked frame (an older build's) that no marked frame follows. - the index's checkpoint, which claims only synced positions (it takes
-
Two frame layouts, one reader (theseus-7nfj, format 6). Frames are marked (magic
THWM, the mark the body's first 8 bytes, the crc over the magic and the body) or unmarked (THWL, every frame before it). The writer writes marked frames only; a log holds unmarked ones until its segments rotate, so the reader tells them apart frame by frame by the magic (wal::Layout), and a magic that rots into the other's fails the crc. A crc-valid frame whose mark is not before its first position is wrong, never torn. Anything that walks frames by hand goes throughLayout(repair.rs) orwal::first_position/wal::read_frame. -
A checkpoint takes the store's
appendinglock alone, so the position it claims is synced and indexed: the writer holds it shared from a batch's first write to its index. Don't take a checkpoint while holding that lock. The periodic one (every 1,000 records) runs on the writer, after it has answered the batch that crossed the mark, so no append's call pays it (theseus-avvb); an append that queues meanwhile waits for it, as it would wherever it ran. A stop's checkpoint (checkpoint_for_close) syncs nothing of its own: redb's close, a durable commit, makes it durable (theseus-02k). Only a durable checkpoint advancesdurable_to, so a durable one after it is never skipped as free. -
The terms and sums are a projection, whole only when marked (theseus-lv2). An open with a
Projectionkeeps each keyed record's terms (the kernel's: an execution's state, …) and numbers (the core's: a session's turns, tokens, and cost, added up per kind) with every append and replay. A checkpoint marks them whole under the projection's name; a writer with no projection (an older build, a tool) moves the checkpoint alone, and the next projected open leaves them tobuild_terms, after serving, never at open. Until they are whole, the store answerslatest_by_terms,count_by_terms, andtotalswithNone, and the reader reads every record. -
The shape is a projection too, whole only when marked (theseus-vm3n.5). With every append and replay, in the transaction that indexes the frame, the index keeps counts (
counts,keycounts,scopecounts,termcounts), a clock per kind (clock) withbytime(a window of time's first position), the records' tags (tagged,pages::tags_of), and each key's birth (born,bybirth). A checkpoint marks them whole underSHAPE(index.shape.3). An open that finds the mark anywhere but at the checkpoint (an older build wrote last) drops only these tables and still reads only the WAL's tail;build_shaperebuilds them after serving, a stretch at a time, andrecountcounts every table in one transaction. Until the build is whole, the counts walk, and a page by tag or time and the newest keys by birth areNone, so their readers scan as before. The build writes with no sync, so it ends with a durable checkpoint of its own, whichbuiltkeeps from being skipped as free: redb writes the build's pages then, after serving, and not in the next stop's close (theseus-celu.16.1). An index with no checkpoint is emptied, all butverified.*, and replayed whole. A change to what these tables hold renames the mark, neverMANIFEST_FORMAT: the index is rebuilt, not read in place. -
The history check starts at the last one's mark (theseus-0dq):
verified.*in the index's meta, written with the next checkpoint. Its frame is checked again first; a frame that no longer checks or holds other positions sends the check back to the log's start, which finds what is wrong. -
Never delete what can't be rebuilt. An index that is not a database is moved aside as
index.redb.bad-<unix ms>, under the file's lock, and rebuilt from the WAL (Item 17). A real database that fails another way is refused: its recovery istheseusd restorefrom the WAL directory. -
A second opener waits for the store's lock up to
LOCK_WAIT(3 s), then fails.
Tests
- The crate's own tests, and
theseus-sim crash-test(a real SIGKILL against a worker process, torn tails, a full disk). Use--restarts 8, so the stores cross checkpoints and most restarts take the tail-only path, and--writers 4, so the kills land among frames the writer commits together. - A test holds the writer by taking
appendingfor writing, as a checkpoint does: what is appended meanwhile queues, and goes in one batch (the_writer_commits_every_queued_frame_with_one_sync). crates/theseus-core/tests/fixtures/store-460a35bis a store an older binary wrote (its README says how it was made). Tests copy it. Never open it in place.crates/theseusd/tests/versions.rsholds the real daemon to the rules, includinga_clean_stop_closes_the_index_and_the_next_start_repairs_nothing.
Traps
index_repaired: true, or a replay, after a clean stop is a bug: something outlived the runtime while holding the store. Since the open makes nothing durable, a run with no durable commit leaves no repair behind, so the replay is the surer sign. Keep statics and detached threads on aWeak.- After
theseus shutdownthe daemon holds the store for 10 to 17 ms more (redb's close). Wait for the process to exit before you read the store's files. - A kill keeps the page cache, so a crash test can't show a missing sync. Test the syncs themselves, as
restore'sDurabledoes (Item 19), and aswal'sa_new_segments_name_is_synced_before_its_first_frame_is_reported_durablecounts the directory syncs.
