Deletion checkpoint persists committed deletes against cold, immutable LWC
rows. It converts the eligible prefix of the in-memory
ColumnDeletionBuffer into persistent ColumnBlockIndex delete metadata and
applies matching secondary-index deletes to DiskTree.
Deletion checkpoint is not an independent root publisher. It runs inside the user-table checkpoint described in Checkpoint, shares that attempt's mutable CoW root and cutoff, and publishes delete metadata together with all companion index state.
Its central durability boundary is deletion_cutoff_ts: cold-row deletes with
commit timestamps below this exclusive cutoff are already covered by durable
table state.
ColumnDeletionBuffer is a concurrent table-level map keyed by RowID. It is
both the write-ownership record for a cold row and the MVCC overlay above
persistent delete metadata.
| Marker state | Meaning |
|---|---|
| Shared transaction status | The delete/update is still owned by a transaction, or its final status is observed through that shared record |
| Compacted commit timestamp | Maintenance has replaced a committed shared status with its immutable CTS |
Foreground cold delete/update, rollback, purge, checkpoint selection, and recovery replay all use this map. Commit becomes visible atomically through the shared transaction status; maintenance may compact the marker later without changing its MVCC meaning.
Foreground claiming distinguishes a foreign preparing owner from an ordinary active owner. The first waiter lazily installs an event on the owner's shared transaction status, releases the deletion-buffer entry guard, waits, checks engine poison, and restarts from authoritative row location and marker state. Point update/delete therefore never consume a stale cold identity after wake. Full-table mutation retains an original-row cursor and owned staged callback output: a pre-callback wait may reload the row and invoke the callback at most once, while a post-callback wait retries the staged action without invoking the callback again.
put_ref() remains the non-waiting maintenance boundary. Page transition,
purge, checkpoint, recovery, and replay callers treat any foreign active marker
as their existing immediate conflict or invariant outcome and never create an
unused prepare listener. Successful failed-precommit rollback removes its cold
marker before prepare completion wakes foreground waiters; fatal cleanup wakes
them only after engine poison is published.
A foreground claim installs undo only for a fresh marker. A same-transaction consumed row does not acquire another independently removable claim. Cleanup validates an active Ref by status allocation while holding the map entry guard, then synchronously keeps or removes that same entry. For transitioned hot rows, page pin, page-state read lock, and row write latch precede the CDB guard. The inverse and exact undo unlink happen inside this guarded reconciliation; no await, I/O, index access, other page acquisition, or user callback occurs there. A surviving same-owner active main predecessor keeps the Ref. Final restoration of a live row removes it. Forward-slot restoration leaves source ownership untouched. Absent, foreign, or committed markers fail before any row mutation.
Deletion checkpoint selects only committed markers below its fixed cutoff. Active markers are ineligible, and a later writer commit receives a later CTS. Removing a rolled-back active marker therefore cannot withdraw a committed delete already selected for checkpoint. Transition installs staged markers once; root publication never recreates a marker that cleanup has removed.
Each ColumnBlockIndex leaf entry stores the owning LWC block's persistent
delete set inline. The set identifies deleted physical row positions; deleting
a row does not change its logical identity or physical membership.
The persisted set is the cold base state and is validated with the owning index entry. Changes become visible only through atomic CoW root publication.
Persistent cold secondary-index entries live in per-index DiskTree roots.
Deleting a cold row therefore has two durable effects:
- add its physical ordinal to the owning LWC block's persistent delete set
- remove or replace the corresponding cold secondary-index entries
Both effects must be present in one table-root publication. A delete cutoff
must never advance past a row whose persistent delete set changed without the
matching DiskTree update.
The in-memory marker is the newest authority when one exists; persistent delete metadata is consulted only when the marker is absent.
For a committed memory marker:
- when its CTS is visible to the reader snapshot, the row is deleted
- when its CTS is newer than the reader snapshot, the marker provides the undo fact that keeps the older row visible
For an uncommitted marker, the owning transaction sees its own delete while other readers preserve snapshot visibility. Competing writers use the marker as the cold-row ownership conflict. A foreground writer may wait only when the owner is preparing; an ordinary active owner remains an immediate conflict. After wake it reclassifies committed, rolled-back, same-owner, replacement, or fatal state instead of assuming that prepare succeeded.
When no memory marker exists, membership in the persisted delete set is final for all active readers: membership means deleted; absence means no cold delete applies.
This ordering is why persistence and in-memory cleanup are separate. A marker may remain necessary for an old snapshot after its delete is present on disk.
One table-checkpoint attempt uses the purge-published GC horizon as its
exclusive cutoff_ts. It reads the mutable root's previous
deletion_cutoff_ts and selects only markers that satisfy all of these rules:
- the marker is committed
previous deletion_cutoff_ts <= cts < cutoff_tsrow_id < pivot_row_id
The lower bound avoids reapplying markers already covered by an earlier checkpoint but retained in memory for MVCC. The upper bound excludes newer commits and deletes just transferred from transitioning row pages. The pivot excludes hot rows whose state is owned by the RowStore image.
Selection is sorted and deduplicated before persistent lookup. If eligible markers exist but the active root has no usable column index or a selected RowID cannot be resolved to an LWC block, checkpoint fails closed and does not advance the cutoff.
Selected RowIDs are resolved and grouped by their persisted LWC block. For each affected block, checkpoint:
- Resolves selected RowIDs to physical row positions.
- Excludes rows already marked deleted in persistent state.
- Reads the old values of the newly deleted rows.
- Derives the corresponding secondary-index changes.
- Merges the new deletions with the persistent set.
A selected RowID inside a block's covered range but absent from its physical membership is an integrity failure.
Old row values must remain decodable until this work publishes. Compaction, vacuum, or page reclamation cannot discard them after a foreground delete but before the table root contains both its delete metadata and companion secondary-index work.
The reconstructed old key determines the exact DiskTree change:
- a unique index removes or replaces the mapping only when the stored owner still matches the deleted old RowID
- a non-unique index removes the exact logical-key and old-RowID entry
The changes accumulate in the same secondary-index sidecar used by data
checkpoint. They are applied to the same mutable table-file fork after all data
and deletion work has been collected. No DiskTree root can be published on
its own.
For every changed LWC block, checkpoint replaces its complete inline deletion set through CoW. The shared block row limit guarantees space for any deletion subset. Physical row identity and values remain intact, including in blocks whose rows are all deleted.
After all patches are represented, the mutable root advances
deletion_cutoff_ts to cutoff_ts. The generic table-checkpoint path then
applies companion secondary roots, rebuilds allocation reachability, and
publishes one atomic root containing:
- the replacement persistent delete metadata
- updated secondary-index
DiskTreeroots - the new exclusive deletion cutoff
- any data-checkpoint state produced by the same attempt
An error before root publication discards the mutable fork. After publication becomes irreversible, a later checkpoint system-transaction failure poisons storage rather than exposing the root as a retryable partial outcome.
An empty selection still proves that no committed cold delete exists in the
scanned timestamp interval. The mutable checkpoint state may therefore advance
deletion_cutoff_ts even when it writes no delete payload.
If some real table-file state changes in the same attempt, the new cutoff is
carried by the published user-table root. If only replay bounds advance, the
checkpoint does not write a metadata-only table root. It upserts the monotonic
table bounds in catalog.table_replay_silent_watermarks; those bounds become
restart proof only after catalog checkpoint folds the row into catalog.mtb.
The silent-watermark publication and fieldwise overlay rules are defined in Checkpoint.
Publishing a delete does not immediately remove its marker from
ColumnDeletionBuffer. Cleanup additionally needs:
- durable delete coverage for the marker
- a snapshot horizon strictly newer than the marker's commit timestamp
Until both proofs hold, the marker may still make a disk-deleted row visible to
an older reader. Secondary MemIndex cleanup has additional row and root
proofs and must not infer safety from deletion-buffer absence alone.
See Garbage Collection for the authoritative cleanup rules.
Recovery loads the persistent delete sets from the selected table root and
reconstructs only the newer tail in memory. For a cold RowID, redo is replayed
into ColumnDeletionBuffer when its CTS is at or above the effective
deletion_cutoff_ts; older delete redo is already covered by persistent state.
Checkpointed silent watermarks may raise the effective cutoff above the value stored in the table root. The complete replay-floor calculation and restart ordering belong to Recovery, avoiding a second recovery algorithm in this document.
Deletion checkpoint keeps three ideas aligned:
- The in-memory marker remains the MVCC authority until cleanup is safe.
- Persistent delete metadata and companion
DiskTreechanges publish through one table root. deletion_cutoff_tsadvances only across a fully examined, durably covered timestamp interval.