Skip to main content

Module chunk

Module chunk 

Source
Expand description

ColumnChunk: differential’s Chunk over ColumnBody updates.

A chunk is a sorted, consolidated run of (D, T, R) updates in the flat columnar layout, in one of two homes:

  • Resident: an Rc-shared ColumnBody on the heap. Fresh input, merge output, and small tails live here.
  • Spilled: the serialized body in the process Pool, with the record count and the first and last data items resident. The pool owns residency from there, with slots under a memory budget and compression and device pageout under pressure, and a body that dies before pressure reaches it is freed without I/O.

Reads of a spilled body are copy-out and scoped to the call that needs them, the contract that lets the pool evict with no reader accounting (see mz_ore::pool).

Spilling happens in Chunk::settle, the trait’s designated commit point: chunks moved to settled output are handed to the pool when spilling is enabled (see set_compute_spill_enabled and set_storage_spill_enabled for how the per-commit destination resolves). Grading is by serialized bytes, the ship size Column already targets, rather than by the record-count TARGET, since record count does not bound bytes for variable-width data.

Chunks whose data is a (key, val) pair additionally implement UnloadChunk, the bulk-read capability: sorted probe keys in, matching updates appended to caller-owned staging, with locate answered from the resident fence metadata so a probe set faults only the chunk bodies it actually touches.

Modules§

metrics
Process-wide work counters for columnar chunk maintenance.

Structs§

AccountedChunkMerger
The chunk merger of AccountedChunkBatcher: differential’s merger with a resident readied side, plus the Merger::allocation figures the mz_arrangement_batcher_*_raw introspection tables are assembled from.
ChunkChunker
A chunker for arrange_core over ColumnChunks: sorts and consolidates raw input columns through a ColumnChunker and wraps its output chunks.
Lz4Codec
The chunk-side ExtentCodec: a little-endian u32 body-length prefix followed by one lz4 block, the framing lz4_flex::block::compress_prepend_size produces. Every chunk consumer passes LZ4_CODEC at insert; the pool itself has no codec opinion.
SpilledBody
A spilled chunk body: the serialized column in the pool, plus the resident metadata every Chunk must answer without fetching. That metadata is the record count, the first and last data items (the fence entries UnloadChunk::locate consults), and the time bounds extract consults to pass frontier-disjoint chunks through without loading them.
UnchunkBuilder
A batch builder over ColumnChunk input that delegates to a builder over ColumnBody input, loading each chunk’s body as it is pushed.

Enums§

ColumnChunk
A sorted, consolidated run of (D, T, R) updates, resident or spilled.

Constants§

COMMIT_BYTES 🔒
The serialized-byte size committed chunks aim for, matching the ship size of the columnar merge machinery.
DEFAULT_COMPRESS_MIN_DEPTH 🔒
The default compression depth floor: fresh (depth 0) bodies spill uncompressed.
READ_SCRATCH 🔒
Reusable staging for call-scoped reads of spilled bodies.
SCRATCH_RETAIN_WORDS 🔒
Scratch capacity retained across reads, in words. A read larger than this releases the buffer afterward, so a thread’s scratch does not ratchet to the largest body it ever carried (heap no pool gauge can see).
SPILL_MIN_BYTES 🔒
Bodies smaller than this stay resident: the pool’s smallest size class is 64 KiB, so spilling below it trades no meaningful memory for slot waste.
SPILL_OVERRIDE 🔒
A thread-scoped pool override, taking precedence over the global enable flag and pool. Lets tests and benches spill through a private pool without touching process-global state.

Statics§

COMPRESS_MIN_DEPTH 🔒
The youngest generational depth whose spilled bodies are compressed. See set_compress_min_depth.
COMPUTE_SPILL_ENABLED 🔒
Compute’s leg of the process spill gate. See set_compute_spill_enabled.
LZ4_CODEC
The Lz4Codec instance chunk consumers pass to Pool::insert_with.
SINK_SPILL_ENABLED 🔒
The gate for bodies spilled through try_spill_ref. See set_sink_spill_enabled.
STORAGE_SPILL_ENABLED 🔒
Storage’s leg of the process spill gate. See set_storage_spill_enabled.

Traits§

ChainState
A builder whose batches carry state derived from a whole chain, computed before the chain’s first push.

Functions§

at_commit_size 🔒
Whether a body is big enough to commit on its own.
chunk_spill_enabled 🔒
Whether ColumnChunks spill: the OR of the compute and storage legs.
codec_for_depth 🔒
The codec a body at depth stores under, identity below the compression floor and lz4 at and past it, paired with whether that codec compresses. One read of the floor, so the pair cannot disagree with itself when the floor moves under a concurrent commit.
compress_min_depth 🔒
The depth floor in effect for this thread’s commits.
cut_records 🔒
Records of a body with len records in bytes that fit in space bytes at the body’s average width, keeping 5% of COMMIT_BYTES as a margin for per-piece framing. Uniform rows cut this way fill their size class and leave one short remainder that settle can coalesce. Uneven widths can still overflow, so callers check the cut piece.
extract_view_into 🔒
Append every update in view whose key matches a probe at or after *probe_index into staging, per the UnloadChunk consume-index protocol: probes strictly below the view’s last key are consumed, a probe equal to it is extracted but left for the next chunk.
resolve_pool 🔒
The thread’s override pool, else the installed pool while enabled.
rr 🔒
Narrow a columnar ref to a shorter lifetime, so refs from different borrows, such as a probe column and a chunk’s own columns, can be compared (the refs are lifetime-invariant).
set_compress_min_depth
Set the youngest generational depth whose spilled bodies are compressed.
set_compute_spill_enabled
Enable or disable chunk spilling on behalf of compute’s arrangement batchers.
set_sink_spill_enabled
Enable or disable spilling of bodies offered through try_spill_ref, which is how compute’s MV sink correction buffer spills.
set_spill_override
Set or unset the pool through which this thread’s chunk spills are routed, taking precedence over the gates and the process pool. None restores the global resolution.
set_storage_spill_enabled
Enable or disable chunk spilling on behalf of storage’s upsert dataflows.
spill_pool 🔒
The pool committed chunks spill to, if any.
spill_serialized 🔒
Serialize a body into a pool slot, writing its serialized form through a cursor over the slot memory. A serialized body’s encoding is one copy. Sizing is exact, so a short or overlong write is a contract violation and panics.
spill_target 🔒
The pool a body of len_bytes spills into while enabled, or None when it stays resident.
try_spill_ref
Spill a serialized copy of body into the process pool, leaving body untouched, or None when the body stays resident.
with_scratch 🔒
Run f with this thread’s read scratch, cleared of any previous use.

Type Aliases§

AccountedChunkBatcher
The ChunkBatcher of a chunk chain, reporting resident bytes to the batcher size logger.