Skip to content

Latest commit

 

History

History
240 lines (179 loc) · 10.9 KB

File metadata and controls

240 lines (179 loc) · 10.9 KB

Deletion Checkpoint

Overview

Deletion checkpoint persists committed deletes against cold, immutable LWC rows. It converts the eligible prefix of the in-memory ColumnDeletionBuffer into persistent ColumnBlockIndex delete metadata and applies matching secondary-index deletes to DiskTree.

Deletion checkpoint is not an independent root publisher. It runs inside the user-table checkpoint described in Checkpoint, shares that attempt's mutable CoW root and cutoff, and publishes delete metadata together with all companion index state.

Its central durability boundary is deletion_cutoff_ts: cold-row deletes with commit timestamps below this exclusive cutoff are already covered by durable table state.

State Model

In-Memory Delete Markers

ColumnDeletionBuffer is a concurrent table-level map keyed by RowID. It is both the write-ownership record for a cold row and the MVCC overlay above persistent delete metadata.

Marker state Meaning
Shared transaction status The delete/update is still owned by a transaction, or its final status is observed through that shared record
Compacted commit timestamp Maintenance has replaced a committed shared status with its immutable CTS

Foreground cold delete/update, rollback, purge, checkpoint selection, and recovery replay all use this map. Commit becomes visible atomically through the shared transaction status; maintenance may compact the marker later without changing its MVCC meaning.

Foreground claiming distinguishes a foreign preparing owner from an ordinary active owner. The first waiter lazily installs an event on the owner's shared transaction status, releases the deletion-buffer entry guard, waits, checks engine poison, and restarts from authoritative row location and marker state. Point update/delete therefore never consume a stale cold identity after wake. Full-table mutation retains an original-row cursor and owned staged callback output: a pre-callback wait may reload the row and invoke the callback at most once, while a post-callback wait retries the staged action without invoking the callback again.

put_ref() remains the non-waiting maintenance boundary. Page transition, purge, checkpoint, recovery, and replay callers treat any foreign active marker as their existing immediate conflict or invariant outcome and never create an unused prepare listener. Successful failed-precommit rollback removes its cold marker before prepare completion wakes foreground waiters; fatal cleanup wakes them only after engine poison is published.

A foreground claim installs undo only for a fresh marker. A same-transaction consumed row does not acquire another independently removable claim. Cleanup validates an active Ref by status allocation while holding the map entry guard, then synchronously keeps or removes that same entry. For transitioned hot rows, page pin, page-state read lock, and row write latch precede the CDB guard. The inverse and exact undo unlink happen inside this guarded reconciliation; no await, I/O, index access, other page acquisition, or user callback occurs there. A surviving same-owner active main predecessor keeps the Ref. Final restoration of a live row removes it. Forward-slot restoration leaves source ownership untouched. Absent, foreign, or committed markers fail before any row mutation.

Deletion checkpoint selects only committed markers below its fixed cutoff. Active markers are ineligible, and a later writer commit receives a later CTS. Removing a rolled-back active marker therefore cannot withdraw a committed delete already selected for checkpoint. Transition installs staged markers once; root publication never recreates a marker that cleanup has removed.

Persistent Delete Metadata

Each ColumnBlockIndex leaf entry stores the owning LWC block's persistent delete set inline. The set identifies deleted physical row positions; deleting a row does not change its logical identity or physical membership.

The persisted set is the cold base state and is validated with the owning index entry. Changes become visible only through atomic CoW root publication.

Secondary-Index State

Persistent cold secondary-index entries live in per-index DiskTree roots. Deleting a cold row therefore has two durable effects:

  • add its physical ordinal to the owning LWC block's persistent delete set
  • remove or replace the corresponding cold secondary-index entries

Both effects must be present in one table-root publication. A delete cutoff must never advance past a row whose persistent delete set changed without the matching DiskTree update.

Read Visibility

The in-memory marker is the newest authority when one exists; persistent delete metadata is consulted only when the marker is absent.

For a committed memory marker:

  • when its CTS is visible to the reader snapshot, the row is deleted
  • when its CTS is newer than the reader snapshot, the marker provides the undo fact that keeps the older row visible

For an uncommitted marker, the owning transaction sees its own delete while other readers preserve snapshot visibility. Competing writers use the marker as the cold-row ownership conflict. A foreground writer may wait only when the owner is preparing; an ordinary active owner remains an immediate conflict. After wake it reclassifies committed, rolled-back, same-owner, replacement, or fatal state instead of assuming that prepare succeeded.

When no memory marker exists, membership in the persisted delete set is final for all active readers: membership means deleted; absence means no cold delete applies.

This ordering is why persistence and in-memory cleanup are separate. A marker may remain necessary for an old snapshot after its delete is present on disk.

Selection Boundary

One table-checkpoint attempt uses the purge-published GC horizon as its exclusive cutoff_ts. It reads the mutable root's previous deletion_cutoff_ts and selects only markers that satisfy all of these rules:

  • the marker is committed
  • previous deletion_cutoff_ts <= cts < cutoff_ts
  • row_id < pivot_row_id

The lower bound avoids reapplying markers already covered by an earlier checkpoint but retained in memory for MVCC. The upper bound excludes newer commits and deletes just transferred from transitioning row pages. The pivot excludes hot rows whose state is owned by the RowStore image.

Selection is sorted and deduplicated before persistent lookup. If eligible markers exist but the active root has no usable column index or a selected RowID cannot be resolved to an LWC block, checkpoint fails closed and does not advance the cutoff.

Block-Grouped Reconstruction

Selected RowIDs are resolved and grouped by their persisted LWC block. For each affected block, checkpoint:

  1. Resolves selected RowIDs to physical row positions.
  2. Excludes rows already marked deleted in persistent state.
  3. Reads the old values of the newly deleted rows.
  4. Derives the corresponding secondary-index changes.
  5. Merges the new deletions with the persistent set.

A selected RowID inside a block's covered range but absent from its physical membership is an integrity failure.

Old row values must remain decodable until this work publishes. Compaction, vacuum, or page reclamation cannot discard them after a foreground delete but before the table root contains both its delete metadata and companion secondary-index work.

Companion Secondary-Index Deletes

The reconstructed old key determines the exact DiskTree change:

  • a unique index removes or replaces the mapping only when the stored owner still matches the deleted old RowID
  • a non-unique index removes the exact logical-key and old-RowID entry

The changes accumulate in the same secondary-index sidecar used by data checkpoint. They are applied to the same mutable table-file fork after all data and deletion work has been collected. No DiskTree root can be published on its own.

CoW Persistence and Cutoff Advancement

For every changed LWC block, checkpoint replaces its complete inline deletion set through CoW. The shared block row limit guarantees space for any deletion subset. Physical row identity and values remain intact, including in blocks whose rows are all deleted.

After all patches are represented, the mutable root advances deletion_cutoff_ts to cutoff_ts. The generic table-checkpoint path then applies companion secondary roots, rebuilds allocation reachability, and publishes one atomic root containing:

  • the replacement persistent delete metadata
  • updated secondary-index DiskTree roots
  • the new exclusive deletion cutoff
  • any data-checkpoint state produced by the same attempt

An error before root publication discards the mutable fork. After publication becomes irreversible, a later checkpoint system-transaction failure poisons storage rather than exposing the root as a retryable partial outcome.

Empty Selection and Silent Progress

An empty selection still proves that no committed cold delete exists in the scanned timestamp interval. The mutable checkpoint state may therefore advance deletion_cutoff_ts even when it writes no delete payload.

If some real table-file state changes in the same attempt, the new cutoff is carried by the published user-table root. If only replay bounds advance, the checkpoint does not write a metadata-only table root. It upserts the monotonic table bounds in catalog.table_replay_silent_watermarks; those bounds become restart proof only after catalog checkpoint folds the row into catalog.mtb.

The silent-watermark publication and fieldwise overlay rules are defined in Checkpoint.

Memory Cleanup

Publishing a delete does not immediately remove its marker from ColumnDeletionBuffer. Cleanup additionally needs:

  • durable delete coverage for the marker
  • a snapshot horizon strictly newer than the marker's commit timestamp

Until both proofs hold, the marker may still make a disk-deleted row visible to an older reader. Secondary MemIndex cleanup has additional row and root proofs and must not infer safety from deletion-buffer absence alone.

See Garbage Collection for the authoritative cleanup rules.

Recovery Boundary

Recovery loads the persistent delete sets from the selected table root and reconstructs only the newer tail in memory. For a cold RowID, redo is replayed into ColumnDeletionBuffer when its CTS is at or above the effective deletion_cutoff_ts; older delete redo is already covered by persistent state.

Checkpointed silent watermarks may raise the effective cutoff above the value stored in the table root. The complete replay-floor calculation and restart ordering belong to Recovery, avoiding a second recovery algorithm in this document.

Summary

Deletion checkpoint keeps three ideas aligned:

  1. The in-memory marker remains the MVCC authority until cleanup is safe.
  2. Persistent delete metadata and companion DiskTree changes publish through one table root.
  3. deletion_cutoff_ts advances only across a fully examined, durably covered timestamp interval.