Skip to content

perf(gfql): answer a numeric datetime searchAny on cuDF from the component index too - #2113

Open
lmeyerov wants to merge 1 commit into
fix/gfql-datetime-index-cache-thrashfrom
perf/gfql-datetime-index-cudf
Open

lmeyerov wants to merge 1 commit into
fix/gfql-datetime-index-cache-thrashfrom
perf/gfql-datetime-index-cudf

Conversation

@lmeyerov

@lmeyerov lmeyerov commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

What

The datetime component index was pandas-only; a numeric searchAny term over a cuDF datetime
column still rendered every row to text on every search. The index needs nothing pandas-specific
— integer components, comparisons, a gather — so it now lives where the column lives: numpy
arrays for a pandas column, cupy arrays on the device for a cuDF one, chosen from the column's
own module. The "cudf" not in type(s).__module__ gate in search_any.py is gone.

cuDF stays UTC-only, for the reason the render is: libcudf's strftime ignores tz_convert,
so no other zone can be named on the GPU. The index declines any other zone with the same
CudfTemporalTzUnsupported the render raises, rather than indexing a false one.

The cache key gains the device: identical bytes on host and GPU are the same instants but not
the same arrays. The cuDF digest is sha256 over a host copy — measured on dgx, a device-side
order-sensitive digest is 30.8 ms but needs a 480 MB powers table, twice the index it protects,
when memory is already the binding constraint; the copy+sha256 is 124.7 ms and needs nothing.

Why this matters: the production table search runs on cuDF frames

The viz's remote table search (GRAPHISTRY_TABLE_QUERY_REMOTE, FEP compute_table_query) works
on cuDF frames, so the pandas index never reached it. Measured on dgx at 30M rows, same frame,
FEP's own match logic lifted verbatim, every term asserted equal:

per keystroke FEP cuDF (GPU render + scan) pygraphistry cuDF before pygraphistry cuDF after
2 (five fields) 861 ms 3,964 ms 133 ms
2024 848 ms 3,935 ms 255 ms
typing sequence 4,259 ms 19,745 ms 783 ms

30× on our own cuDF path; 5.4× faster than the production FEP path on FEP's frames. The
residual ~125 ms is the host-copy digest; a caller-supplied cache key (FEP has dataset ids) would
take a keystroke to ~3–70 ms.

Typing

compute/typing.py's ArrayLike/ArrayNamespace protocols gain the members the kernel calls
(equal, logical_or, logical_and, take, int8/int16; __mod__, all, any, copy),
so the module-agnostic code is typed against the same protocols gfql/index/engine_arrays uses.
The type-hygiene guard rejected cast(...) in favour of engine-agnostic annotations with
localized type: ignores — followed.

Tests

The cuDF class (TestOnCuDF, 9 tests) is skipped without a GPU and validated on dgx: 80 passed
— exactness against render_datetime_cudf over every shape and term, the route pinned on a cuDF
frame, the non-UTC decline, and that host and device columns with the same bytes do not share an
entry. CPU suites unchanged. Changed-line coverage vs the stack top 52/61 = 85%, no pragmas:
the nine uncovered lines are the GPU-only arms, which dgx covers.

Stack

PR 4 of 4, base fix/gfql-datetime-index-cache-thrash (#2112). Lands after #2110 → #2111 →
#2112; GitHub retargets the base on each merge.

🤖 Generated with Claude Code

https://claude.ai/code/session_017ropeBMLJUuy6ViYwy15ud

Check count

This PR will show 68 checks rather than the 70 on #2110: codeql-analysis.yml runs on pull_request only for branches: [master], so a stacked PR whose base is another branch skips it until the base retargets to master after the PRs below it merge. The 68 are the full CI Tests matrix.

…onent index too

The index was pandas-only; a cuDF datetime column still rendered every row to text on every
search. The index needs nothing pandas-specific -- integer components, comparisons, a gather --
so it now lives where the column lives: numpy arrays for a pandas column, cupy arrays on the
device for a cuDF one, chosen by the column's own module.

cuDF stays UTC-only, for the reason the render is: libcudf's strftime ignores tz_convert, so
no other zone can be named on the GPU. The index declines any other zone with the same
CudfTemporalTzUnsupported the render raises, rather than indexing a false one.

The cache key gains the device. Identical bytes on the host and on the GPU are the same
instants, but their indexes are arrays on different devices and cannot be shared. The cuDF
digest is taken on a host copy: a device-side digest either copies anyway or needs a powers
table larger than the index it protects.

The typing protocols in compute/typing gain the members the kernel calls (equal, logical_or,
logical_and, take, int8/int16; __mod__, all, any, copy on arrays), so the module-agnostic code
is typed against the same ArrayLike/ArrayNamespace the gfql index engine already uses.

Tests: the cuDF class is skipped without a GPU and validated on dgx -- exactness against
render_datetime_cudf over every shape and term, the route pinned on a cuDF frame, the non-UTC
decline, and that host and device columns with the same bytes do not share an entry.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017ropeBMLJUuy6ViYwy15ud

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant