Remove pgrust: the entry is cheating (per-query kernels and result caching) - #2165
Merged
Merged
Conversation
…s their results The v0.4-preview binary embeds one hand-written kernel per ClickBench query (crates/gravity/backend/irprobe/src/q00.rs ... q36.rs), plan matchers named after the benchmark queries, and a cross-query cache of filter verdicts that makes the hot runs replay the answer of the previous run without touching data. Trivially perturbed queries such as SELECT MIN(URL) FROM hits are refused with an error. See the pull request description for the details.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This removes the
pgrustentry (added in #1163, updated to v0.4-preview in #2106) from the benchmark. The results it reports are produced by machinery written specifically for the 43 ClickBench queries, and its hot runs replay cached answers of the previous run. Both violate the rules: query result caches and caches near the end of the pipeline must be disabled, and the entry is taggedtuned: nowhile it is tuned for these exact queries.Everything below was verified against the published
v0.4-previewrelease binary (pgrust.com/downloads/v0.4-preview/, the oneinstalldownloads) and the publicv0.3source (github.com/malisper/pgrust). The v0.4-preview source is not published; the binary identifies itself aspgrust 0.3-beta.One hand-written kernel per benchmark query
The aarch64 release binary embeds these source paths:
The numbers are ClickBench query numbers. The plan matchers that dispatch to these kernels emit refusal messages named after the benchmark, for example:
There is also an environment knob literally named
PGRUST_GRAVITY_Q20_VARIANT.The public v0.3 source confirms the design:
crates/backend/executor/sqe/rig/plans/clickbench.ronis described in its header as "hand-lowered physical plans for all 43 ClickBench queries", andqueries-clickbench.sql(the verbatim benchmark SQL) is compiled into the executor crate withinclude_str!. That file classifies Q0, Q1, Q2, Q3, Q6, Q19 and Q29 asMetadataAnswer(answered from per-part statistics, no scan), notes that Q19's "answer lines are reconstructed from the predicate constant", that Q29 is an algebraic rewriteSUM(x+k) = SUM(x) + k*COUNTfollowed by a metadata answer, and that Q18 depends on a persisted derived column "with baked bit widths 26/6/20". Q28's regular expression is a one-entry hardcoded table (HOST_EXTRACT_PATTERN = ^https?://(?:www\.)?([^/]+)/.*$), and a planner comment reads "The 43's only HAVING shape: COUNT(*) > k".The hot runs are a result cache
For Q20, Q21 and Q23 the plan notes say: "hot rep = pure fold over cached verdicts — ZERO columns touched; the selection is the answer".
EXPLAINon the release binary reportsSQE: engine (family: windowreplay)for Q20 andfamily: metaanswerfor Q1 with the scan markednever executed.Measured locally on a 96-core Graviton (Neoverse-V2) machine, using the entry's own scripts and configuration, with a true cold cycle (stop, drop caches, start) like the driver does:
URL LIKE '%google%''%rambler%': 283 ms, then 12 ms, 4 msGROUP BY UserID, SearchPhrase LIMIT 10GROUP BY URL ORDER BY c DESC LIMIT 10A second restart brought Q20 back to 11.2 s, then 3 ms. The engine's real warm speed for a substring scan is a few hundred milliseconds, which is about what ClickHouse does on the same hardware. The published hot numbers (7 to 23 ms for Q20) are the replay of the previous run's verdicts.
The v0.2 submission (#1163) disclosed a similar cache (
pgrust.condition_cache = on) and offered to switch it off. The v0.4 update removed that README and the setting. No configuration disables the cache now: withpgrust.gravity_entry = offandpgrust.condition_cache = off, a repeatedLIKEstill returns in 3 to 4 ms.Queries outside the 43 shapes are refused
The engine has no general execution path for the columnar table. Trivially perturbed benchmark queries fail with
ERROR: sqe: statement over columnar relation "hits" is not served, and the error detail says "The sqe engine serves columnar shapes with no fallback; an unserved shape fails loudly instead of running slowly". 20 of 67 near-benchmark probes and 3 of 49 perturbed variants were refused, including:INSERT,UPDATEandDELETEare not supported on the table either (pgrcolumnar2 does not support INSERT).For the record
The answers to the 43 queries themselves are correct: compared against
clickhouse localover the samehits.parquet, 30 match exactly and the remaining differences are valid tie-breaks onLIMITqueries (verified individually). The problem is not wrong results; it is that the results measure a lookup table of the benchmark, not a database.Related: #1163
Related: #2106