Layer 01: the KV abstraction
The ordered key-value and transaction contract beneath Rad's relational engine.
The KV layer is Rad's storage boundary. The catalog, planner, and executor depend on this contract rather than a storage backend directly.
The layer deals in byte keys, byte values, ordered ranges, and optimistic transactions. It knows nothing about tables or rows. Higher layers create that meaning through key layouts and typed encodings.
Store and iterator contract
The storage vocabulary has four operations:
| Operation | Contract |
|---|---|
| Point read | Return one value or an explicit missing result |
| Write | Replace the value at one key |
| Delete | Remove one key; deleting a missing key succeeds |
| Scan | Visit a half-open key range in ascending or descending lexicographic byte order |
An omitted scan bound is unbounded. Prefix scans use the smallest possible end key after every key beginning with that prefix. This makes a prefix range fit the same ordered-scan contract without adding a separate storage operation.
Iterator key and value buffers remain valid only until the next step. Data that must survive another step or an interleaved read is copied first. Transaction iterators close before commit.
The contract does not specify what an already-open iterator sees after a write to the same transaction. Upper layers take a new scan after such writes.
Transactions
Transactions read a stable snapshot overlaid with their own buffered writes. Commit publishes every buffered write atomically. Rollback discards them.
| Isolation level | Conflict detection | Intended use |
|---|---|---|
Snapshot | Write-write conflicts | Read-only execution and cases where write skew is acceptable |
SerializableSnapshot | Write-write and read-write conflicts, including requested scan ranges | Mutations, catalog changes, and PIR programs with effects |
Serializable range reads protect their requested bounds, even when the scan returns no keys. A concurrent insert into that range conflicts at commit. Rad relies on this for unique-index and foreign-key checks without taking locks.
Conflicts return to the owner of the complete transaction, which may retry from the start. The KV layer does not replay the transaction itself.
Positions
An opaque data position identifies the backend state at which a transaction's snapshot began.
The SlateDB adapter encodes sequence numbers as decimal text, but this is an adapter detail. Higher layers may persist and compare token identity. They must not parse, order, increment, or manufacture a position.
The current API cannot begin a transaction at an old position, reopen a historical snapshot, or tell a backend to retain and later release one. Schema transitions record positions as provenance and barriers; workers do not use them as read-at-position handles.
Execution snapshots
The standalone LIR read path begins one Snapshot transaction, then binds,
plans, admits catalog dependencies, and reads through that transaction's stable
view. It rolls the transaction back when the result is complete because there
are no writes to publish.
PIR execution has two storage phases. Preflight runs against a
SerializableSnapshot transaction and rolls it back. Execution then binds the
program again against a fresh transaction. A read-only program uses Snapshot
and rolls back after producing its result. A program containing any data or
catalog effect uses SerializableSnapshot and commits the complete program.
Catalog dependency admission happens inside the execution transaction. The admitted fence reads remain part of the backend conflict set through commit.
Object storage selection
Slate can use memory, file, or S3 object storage. RAD_STORAGE selects the
implementation. RAD_STORAGE_PATH sets the local database directory or the
object prefix. One write process owns the database. Additional read processes
can follow published checkpoints.
SlateDB adapter
The SlateDB adapter maps snapshot and serializable isolation to SlateDB's corresponding transaction modes. Transaction conflicts are normalised into Rad's conflict category.
The adapter waits for each transaction WAL write to reach object storage before
it returns success. Set --slate-commit-durability memory or
RAD_SLATE_COMMIT_DURABILITY=memory to return after SlateDB accepts the write
in process memory. A process failure can lose these acknowledged writes. An
orderly process stop flushes them to object storage. A storage error can occur
after Rad returns success for an unflushed memory commit.
The adapter checks cancellation before a storage call, but a scan or commit already in progress cannot currently be interrupted.
Rad opens one SlateDB writer process per database. The adapter supports file
and memory object-store URLs; tests exercise SlateDB itself on memory:///
rather than a separate fake backend.
Slate tuning
Rad uses a 128 MiB Foyer cache for decoded blocks and metadata. It does not add throughput scan blocks to this cache by default. This protects point-read cache entries from large scans. A throughput scan reads ahead by 256 KiB and uses at most four fetch tasks.
The local raw object cache is disabled until
--slate-object-cache-path is set. The cache has a 16 GiB capacity and 4 MiB
parts after it is enabled. Flush, compaction, and startup preload do not fill it
by default. Use the slate-object-cache-* options to change this policy.
The remaining slate-* options control WAL flush frequency, L0 SST creation
and backpressure, Bloom filters, and SST block size. The command reference
states each default. Use these controls only when workload and Slate metrics
show a specific cache, read, or write limit.
See the Slate documentation for cache design, Foyer cache operation, and storage tuning.
Ordered encoding
The key encoding makes semantic values sort correctly as bytes. Encodings are tagged and self-delimiting, so several values can be concatenated into a tuple without separators.
| Type | Ordering detail |
|---|---|
NULL | Sorts before every non-null tagged value |
bool | Tagged deterministic order |
int64 | Sign bit transformed, then big-endian |
float64 | Bit transform preserving numeric order; callers reject NaN |
text | NUL escaped and terminated while preserving byte order |
Tags make mixed types deterministic, not numerically interchangeable. Every stored column position must use one consistent type.
Executable evidence
The backend conformance suite covers snapshot stability, own writes and
deletes, write conflicts, point-read conflicts, phantom ranges, rollback
lifecycle, and the write-skew distinction between the two isolation levels.
The SlateDB adapter runs that suite directly. Encoding tests add round-trip and
random byte-order agreement checks, including embedded NUL and 0xFF cases.