How tsoracle is laid out internally — the crate boundaries, the deliberate sync/async split, the trait contracts that hold the layers together. Read this before The Allocator and Key Subsystems: those chapters assume you've seen the layering choices explained here.
tsoracle is a Cargo workspace of seven crates, layered so the algorithm itself stays sync and runtime-neutral and only the server crate carries async:
| Crate | Role |
|---|---|
tsoracle-proto |
gRPC service & message definitions, generated by tonic-prost-build. Stable wire contract. |
tsoracle-core |
The window allocator state machine, Clock trait, Epoch type, Timestamp packing, and the dense-sequence keying/validation types (SeqKey, SeqGrant). Pure sync. No tokio, no I/O. |
tsoracle-consensus |
The ConsensusDriver trait (timestamp persistence plus the optional advance_dense/load_dense_seq methods) and shared types. Where async first appears, but only as a method signature — no runtime is assumed. |
tsoracle-driver-file |
Single-node, fsync-backed ConsensusDriver. The default driver. |
tsoracle-server |
Tonic service, leader-watch task, failover fence. The only crate that depends on tokio directly. |
tsoracle-client |
gRPC client with leader discovery and request coalescing. |
tsoracle-bin |
The tsoracle CLI. |
The deliberate sync/async split is the most important shape: the algorithm crate (tsoracle-core) has no await and no Future. Window extension is exposed as a two-call API — prepare_window_extension (sync, computes the target) and commit_window_extension (sync, applies the persisted value) — so the single await lives in the server. This is what lets the core be property-tested in microseconds with no runtime, and is also why ConsensusDriver lives in its own crate: putting it in core would force core to depend on async_trait and futures::Stream; putting it in server would mean external driver crates would have to depend on server. The standalone trait crate breaks the dependency cycle.
The GetSeq path rides the same layering. The service validates the request against tsoracle-core's keying types (SeqKey/SeqGrant) — pure sync, no counter state — then awaits ConsensusDriver::advance_dense for the durable per-key fetch-add. No window-allocator state is involved; the dense counter is a sibling of the high-water, not part of the timestamp window (see The Allocator → Dense sequences).
The Clock trait's now_ms is advisory. It tells the allocator how to advance physical time toward the wall clock. The allocator's monotonicity does not depend on clock correctness — a clock jumping backward cannot cause regression because the persisted high-water always wins. A clock pinned far in the past stalls new windows until wall time catches up past the persisted bound. Implementing a custom Clock (NTP-disciplined, HLC-extended, externally driven) does not require any changes to the allocator's correctness reasoning.
An Epoch is an opaque monotonic identifier chosen by the consensus driver. The library treats it as a u64 with no internal meaning: drivers typically map it to the consensus layer's term, lease generation, or election number. tsoracle's only requirement is that epochs are non-decreasing across a driver instance's lifetime, and that an epoch passed to persist_high_water matches the epoch under which the calling allocator believes it is leader.
Fencing semantics rely on the epoch:
- The leader-watch pipeline records the epoch on every transition into
Leader { epoch }. The allocator stores it. - Every steady-state extension passes the current epoch to
persist_high_water(at_least, epoch). A driver that accepts the call only when it is the current leader at that epoch (the contract of persist_high_water) makes stale-leader writes impossible. - The library never sees raw consensus node IDs; that mapping (node ID ↔ tsoracle service address, term ↔ epoch) is the driver's job.
Timestamp(u64) is packed as physical_ms << 18 | logical:
- The 46-bit physical field is milliseconds since Unix epoch (
PHYSICAL_MS_MAX = 2^46 - 1, exhausted in year 4199). - The 18-bit logical field is a counter that increments within a single millisecond (
LOGICAL_MAX = 2^18 - 1 = 262 143), and resets to zero each timephysical_msadvances.
The split is calibrated so that 262 144 timestamps per millisecond is the per-instance ceiling — beyond that, the allocator returns WindowExhausted to the next request and the millisecond must advance before serving resumes. At 100 000 timestamps/sec sustained the logical field is barely used (~100 per ms) and physical_ms increases predictably with wall time.
Timestamp::pack(physical_ms, logical) constructs a packed value and panics if either field is outside the 46/18-bit layout. Use Timestamp::try_pack at trust boundaries where invalid input should become a recoverable error. physical_ms() and logical() are the accessors. The Ord impl on Timestamp is the lexicographic comparison of (physical_ms, logical), which agrees with numeric ordering on the packed u64 — important because callers often store the raw u64 in databases and need the durable ordering to match the in-memory ordering.
A different split is possible but not parameterized in the type. If you need more physical headroom (past year 4199) or more logical capacity (past 262 143 per ms), fork the type.