AOF / TsavoriteLog record layout
This document describes the on-log byte layout of non-chunked and chunked AOF records, and how a large key/value/object is split across multiple AOF entries and reassembled on replay.
Keep this document in sync with the code. If any of the following change, update the layouts and diagrams below:
libs/server/AOF/AofHeader.cs— the AOF header hierarchy and theflagsbitfield.libs/server/AOF/AofChunkHeader.cs— the per-chunk framing header.libs/server/AOF/GarnetLog.cs—Enqueue/EnqueueSpanChunked/EnqueueObjectChunked/IsChunkable.libs/storage/Tsavorite/cs/src/core/TsavoriteLog/TsavoriteLog.Chunked.cs— the chunk writer (WriteOneRecord).libs/server/AOF/AofChunkedRecordReader.cs/AofProcessor.ChunkReplay.cs— the reader / replay side.libs/storage/Tsavorite/cs/src/core/Allocator/ObjectSerialization/ChunkedRecordConstants.cs— the continuation flag.
1. Overview
Garnet's AOF is a TsavoriteLog. Each mutating operation (SET, HSET, an object upsert, an RMW,
DELETE, a transaction marker, …) is appended as one or more AOF entries.
- A small operation is written as a single, non-chunked entry (§3–§4).
- An operation whose key + value + input together exceed
TsavoriteLog.MinPartialAllocSizeis written as a chunked record: a run of AOF entries that the reader reassembles (GarnetLog.IsChunkable) (§6).
Chunking lets a value that is larger than an AOF page — even larger than 2 GB — be written and replayed without ever materializing the whole serialized value contiguously. The same chunked-record machinery is reused by cluster migration and diskless replication; §5 maps all of those paths before the byte-level detail.
The AOF stores an operation image (
opType+ key + value/input), which is not the same as the serializedDiskLogRecordimage shipped by cluster migration / replication. For that format see the companion doc, Migration / Replication record layout.
2. TsavoriteLog entry framing
Every AOF entry — chunked or not — is a single TsavoriteLog entry:
+------------------------------+------------------------+-----------------------------+
| entry-length prefix | AOF header (variant) | body |
| (TsavoriteLog headerSize) | §3 / §6 | §4 / §6 |
+------------------------------+------------------------+-----------------------------+
- The entry-length prefix is written by
TsavoriteLog(SetHeader); it frames the entry on the page and is not part of the AOF header. - The AOF header variant is selected by the log topology (single vs. sharded, per-key vs. coordinated).
3. Non-chunked AOF headers
The header hierarchy is defined in AofHeader.cs. All offsets/sizes are in bytes.
AofHeader — 16 B (base, single-log per-key entries)
off 0 1 2 3 4 12 16
+--------+--------+--------+----------+-------------------------+-----------+
|version | flags | opType | proc/db | storeVersion | sessionID |
| u8 | u8 | u8 | id u8 | i64 | i32 |
+--------+--------+--------+----------+-------------------------+-----------+
flags is a bitfield:
bit 7 6 5 4 3 2 1 0
. . . . | +---+---+--- AofHeaderTypeMask (0b0111)
\___________/ | | \_______ base type (bits 0-1): Basic=0, Sharded=1,
unused | | SingleLogTxn=2, ShardedLogTxn=3
| +----------- ChunkedRecordFlag (0b0100): set on chunk entries (bit 2)
+--------------- UnsafeTruncateLogFlag (0b1000): FLUSH (bit 3)
A chunked header type is simply its non-chunked counterpart with bit 2 set: BasicChunkHeader = 4,
ShardedChunkHeader = 5.
Other non-chunked headers
| Header | Size | = base + extra | Used for |
|---|---|---|---|
AofHeader | 16 B | — | Single physical log, per-key entries |
AofShardedHeader | 24 B | AofHeader + sequenceNumber (i64) | Multi-physical-log (sharded) per-key entries |
AofSingleLogTransactionHeader | 50 B | AofHeader + participantCount (i16) + replayTaskAccessVector (32 B) | Coordinated ops (txn/checkpoint/flush), single physical log + multi-replay |
AofShardedLogTransactionHeader | 58 B | AofShardedHeader + participantCount (i16) + replayTaskAccessVector (32 B) | Coordinated ops, sharded |
Selection logic (GarnetLog.Enqueue):
| Topology | Per-key entry | Coordinated / broadcast entry |
|---|---|---|
| Single log (1 physical, 1 replay task) | AofHeader | AofHeader |
| Single physical log, multi-replay | AofHeader | AofSingleLogTransactionHeader |
| Multi physical log, multi-replay | AofShardedHeader | AofShardedLogTransactionHeader |
4. Non-chunked entry body
For a per-key operation, the body that follows the header is the operation's key, then (as required by opType) the
value and/or the serialized input, laid out by TsavoriteLog.Enqueue(header, key, value, ref input, …):
+----------------------+--------------------------+--------------------------+---------------------+
| AOF header (§3) | key | value (Upsert shapes) | input (RMW / |
| | [i32 len][key bytes] | [i32 len][value bytes] | Upsert-with-input) |
| | (SpanByte framing) | (SpanByte framing) | [raw serialized] |
+----------------------+--------------------------+--------------------------+---------------------+
- key and value use SpanByte framing: a 4-byte little-endian length prefix followed by the bytes
(
ReadOnlySpan<byte>.SerializeTo,TotalSize = 4 + Length). - input is the raw serialized input (
IStoreInput.CopyTo,SerializedLengthbytes; no separate length prefix). - Which components are present depends on the op shape: e.g. DELETE is key-only (the value slot is unused on replay); a plain upsert is key + value; an RMW / upsert-with-input carries the input.
5. Chunk management at a glance
Chunking appears in three subsystems — the AOF (write and replay), cluster migration (send and receive), and diskless replication (send only) — plus disk-based replication, which is listed because it is routinely assumed to use this machinery and does not. They share two building blocks: a ring that bounds how much serialized data is held while producing, and an accumulator that collects arriving pieces until a record is complete. Each path combines them differently, and two use only one of the two.
| Path | Producer side | Carrier | Consumer side |
|---|---|---|---|
| AOF write | object value streamed through a 4 MB ring, rented per write from the log's SectorAlignedBufferPool; one AOF entry per drain. A span value needs no ring — it is handed over whole. | AOF entries in the log | — |
| AOF replay | — | AOF entries read by the scan iterator | ChunkedAccumulator: key / span value / input into buffers pre-sized from the chunk header; object value into a PooledChunkList, then deserialized from a ReadOnlySequence |
| Migration send | object value serialized in-epoch through a 4 MB ring owned by the accumulator, draining into a PooledChunkList | one LogRecord if the record fits a send buffer, else ChunkedLogRecord chunks | — |
| Migration receive | — | network commands (a record's chunks may span several) | ChunkedRecordReassembler: inline portion into a contiguous buffer, overflow key/value into a pre-sized OverflowByteArray, object value into a PooledChunkList |
| Diskless replication send | object value streamed in-epoch through a ring sized just under the send buffer — no accumulator; each drain goes straight to the network | replication stream | replica replays as AOF |
| Disk-based replication | not record chunking at all — see below | checkpoint file byte ranges | replica writes files to disk, recovers, then replays AOF |
The two buffer roles are easy to confuse because both are "pooled chunk memory", but they answer to different constraints:
| Ring (producer) | Accumulator (consumer) | |
|---|---|---|
| Holds | serialized bytes not yet drained | pieces already received |
| Size governs | how much value data one carrier chunk holds — so it sets the number of chunks emitted | how many segments the resulting ReadOnlySequence has |
| Affects the wire/log format | yes | no |
| Lower bound | MinPartialAllocSize, or page-tail packing stops working (AOF only) | none |
Why migration accumulates but diskless replication does not. Both must serialize while holding the store epoch,
because a migrating or replicating key is not locked and its value may be updated concurrently. The difference is how
they send. Replication sends synchronously (BlockingWait), so it can stream each drained chunk to the network
without ever leaving the epoch, and never materializes an object value whole. Migration sends asynchronously, and
the store epoch must not be held across an await — so it snapshots the record's pieces in-epoch into the accumulator
and assembles and sends them out of epoch.
Diskless vs disk-based replication. Only diskless replication uses the machinery in this document. It iterates
the live store and streams records, so it needs the chunked-record format for values too large to send whole.
Disk-based replication does something categorically different: the primary takes a checkpoint to disk and ships byte
ranges of the resulting files (CheckpointFileType.STORE_HLOG, STORE_SNAPSHOT, …) via FileDataSource, which
batches through a SectorAlignedBufferPool. Those batches are file offsets, not records — there are no chunk headers,
no component ordering, and no reassembly of a logical record. The replica writes the files, recovers from them, and
only then replays the AOF, at which point the AOF replay row above applies.
Why an arriving object value is accumulated rather than deserialized as it arrives — on both AOF replay and migration receive — is covered in §6.2; the short version is that the records that would benefit are exactly the ones it cannot be applied to.
The rest of this document details the AOF rows. For the migration and replication rows — the DiskLogRecord image,
which component goes where on the wire, and a per-path accounting of copies — see the companion doc,
Migration / Replication record layout.
6. Chunked AOF records
When GarnetLog.IsChunkable(key, value, input) is true (key.TotalSize + value.TotalSize + inputSerializedLength > TsavoriteLog.MinPartialAllocSize), the operation is written as a run of chunk entries by EnqueueSpanChunked
(span key/value) or EnqueueObjectChunked (streamed object value).
The write-side objects — the ChunkWriteState and, for object values, the ChunkedObjectSerializer with its ring buffer
and stream — are cached per thread and rebound per record, so a chunked write allocates nothing on the steady-state path.
The ring is rented per write from the log's SectorAlignedBufferPool and sized to
TsavoriteLog.ChunkedObjectRingBufferSize (IStreamBuffer.BufferSize, 4 MB) rather than from the value. Two things follow
from that size, so it is not free to change:
- Because a chunk entry is allocated per drain, the ring also bounds how much value data one entry carries — a smaller ring does not merely hold fewer bytes at once, it multiplies the number of entries written.
- It sits at or above
MinPartialAllocSize, which page-tail packing requires: an allocation is only split across a page boundary when both halves reach that size, so a ring below 1 MB would silently stop chunk entries filling a page tail.
Pool blocks are pinned arrays that the pool reuses, so the size costs no repeated large-object-heap allocation, and renting per write means a cached serializer holds no buffer (and roots no value object) between writes.
6.1 Chunk headers
Each chunk entry uses a chunked header: a normal header immediately followed by an AofChunkHeader.
AofChunkHeader — 28 B (AofChunkHeader.cs):
off 0 4 8 12 20 28
+------------+------------+------------+--------------------+--------------------+
| overflow | overflow | input | objectId | keyHash |
| KeyLength | ValueLength| Length | (u64) | (i64) |
| u32 | u32 | u32 | | |
+------------+------------+------------+--------------------+--------------------+
overflowKeyLength/overflowValueLength/inputLength— the full length of each component, known up front, so the reader pre-allocates one buffer per component.overflowValueLengthis left 0 for a streamed object value (its length is not known up front; the reader accumulates it into pooled buffers instead).objectId— the identifier that groups a record's chunks: the logicalAddress of the record's first chunk, written identically on every chunk. It is the only field patched per-chunk at write time.keyHash—GarnetLog.HASH(key), identical on every chunk; used to route all of a record's chunks to the same replay task during parallel/sharded replay (the chunks cannot expose the key directly — it is itself split across chunk data).
Chunked header variants:
| Header | Size | = |
|---|---|---|
AofBasicChunkHeader | 44 B | AofHeader (16) + AofChunkHeader (28) |
AofShardedChunkHeader | 52 B | AofShardedHeader (24) + AofChunkHeader (28) |
6.2 Chunk entry body: packed component segments
The components are written in the fixed order Key → Value → Input, packed: a single chunk entry holds one
[i32 prefix][data] segment per component that (partly) fits, in order — so key + value + input can share an entry.
A component too large to fit is split: its last segment in an entry sets the high bit of the prefix (the continuation
flag), and it resumes in the next entry.
chunk entry body = [ prefix | seg-data ] [ prefix | seg-data ] ... (bounded by the entry length)
prefix (i32):
bit 31 = ChunkedRecordConstants.ContinuationFlag (1 = more of this component follows)
bits 30..0 = this segment's data length
- The reader (
AofChunkedRecordReader.ReadChunk) reads every segment in an entry, and advances Key → Value → Input each time it sees a prefix whose continuation flag is clear (the current component is complete). - A length prefix is written whole or not at all: if fewer than
sizeof(int)bytes remain in the entry, the prefix is deferred to the start of the next chunk entry — a prefix is never split across an entry boundary, on both the write (WriteOneRecord) and read side. - A span (overflow) value is copied into its pre-sized buffer; a streamed object value is accumulated into a
PooledChunkListand exposed as aReadOnlySequence<byte>(ChunkedAccumulator.GetValueSequence) for streaming deserialize with no giant contiguous copy — this is what lets an object value exceed 2 GB. The list packs arriving segments into uniform pooled buffers, so transport framing does not dictate the allocation shape.
Why the value is accumulated and not deserialized as it arrives. Accumulating means the serialized bytes are held
alongside the materialized object, and deserializing incrementally would avoid that — but it cannot be applied to the
records that would benefit. Whole-object upserts come from RENAME, which wraps itself in an internal transaction, so
replay is inside a transaction when the chunks arrive and the record is buffered until commit rather than dispatched;
a buffered record holds its value in whichever form it was accumulated, and for a collection the materialized object is
several times the size of the bytes. Deserialization is also synchronous, so incremental deserialization needs a second
thread fed by the replay thread, which cannot block where that thread holds the log epoch (the bulk-consume path runs
the consumer under it for entries resident in the log buffer) or is itself the thread that must deliver the remaining
chunks (replication replay, fed from the network). Between them these exclude every path a chunked object value
currently arrives on.
6.3 A large object across chunk entries
logical record (opType = ObjectStoreUpsert, key K, object value V, |V| >> page)
entry 0 entry 1 ... entry N
+---------------------+ +---------------------+ +---------------------+
| AofBasicChunkHeader | | AofBasicChunkHeader | | AofBasicChunkHeader |
| objectId = addr0 | | objectId = addr0 | | objectId = addr0 |
+---------------------+ +---------------------+ +---------------------+
| [len|+] key seg | | [len|+] value seg | | [len ] value seg | <- last: flag clear
| [len|+] value seg | | [len|+] value seg | | [len ] input seg | <- (if any)
+---------------------+ +---------------------+ +---------------------+
page tail ---------------> next page (page-tail packing via AllocateBlockPartial)
every entry carries objectId = addr0 (the first chunk's logicalAddress) so the reader groups them;
the value's continuation flag stays set until the object serializer's final (isComplete) drain.
6.4 Write and replay flow
7. Call sequence (code paths)
The AOF logs an operation (opType + key + value/input); a chunkable op is split across entries on write and
reassembled on replay. Indentation = call depth; a multi-step flow may sit on one line (a → b → c), and italic
sub-items are terse notes.
Write — a chunkable mutating op logs to the AOF:
GarnetLog.Enqueue(opType, key, value, input)GarnetLog.cs; chunks whenkey + value + input > TsavoriteLog.MinPartialAllocSize(IsChunkable)EnqueueSpanChunked(...)— or —EnqueueObjectChunked(...)- span key/value op vs object-value op
TsavoriteLog.EnqueueChunkedObject(header, key, input, objectSerializer, value)TsavoriteLog.Chunked.csChunkedObjectSerializer.Serialize()→TsavoriteLog.Consume(first, second, key, input)→WriteOneRecord(...)- stream the object value through a bounded ring; each drain packs the components — key, then value, then input — as
[i32 prefix|cont][data]segments into one page-tail chunk entry (AofBasicChunkHeader/AofShardedChunkHeader), splitting a component across entries as needed - carrier: the AOF chunk entries themselves — no write-side accumulator; segments go straight to the log → AOF pages
- stream the object value through a bounded ring; each drain packs the components — key, then value, then input — as
Replay — read the chunk entries back and apply to the store:
AofProcessor.ProcessAofRecordInternal(ptr, length)AofProcessor.cs; per scanned entryAofChunkedRecordReader.ReadChunk(ptr, length, out acc)→AppendChunk(...)- accumulate this entry's segments into the record's
ChunkedAccumulator; true once every component has arrived - populates
ChunkedAccumulator.key/.value(span value) /.input; an object value's chunks append to.valueChunks(PooledChunkList)
- accumulate this entry's segments into the record's
ProcessAofRecordInternal(acc)→ReplayOp(acc)AofProcessor.ChunkReplay.cs; record complete → dispatch by opTypestringContext.Upsert(key, input, value)StoreUpsert: inline / overflow string value → store
ObjectStoreUpsert(acc, ...)→GarnetObjectSerializer.Deserialize(acc.GetValueSequence())→objectContext.Upsert(key, valueObject)- object value: stream-deserialize from the
ReadOnlySequenceover.valueChunks(no giant contiguous copy; supports values over 2 GB) → store
- object value: stream-deserialize from the