Add Apache Kudu - #2262
Add Apache Kudu#2262KazukiKandaKK wants to merge 10 commits into
Conversation
|
Results for Logs:
|
|
The run of Logs:
|
|
Results for Logs:
|
|
Results for Logs:
|
|
Results for Logs:
|
|
The run of Logs:
|
|
Sometimes it fails to load the dataset:
on c6a.metal. |
The bucket-UPSERT loop piped SQL into impala-shell's stdin mode, which exits 0 on EOF even when a query fails, masking OOM/syntax errors behind "row count mismatch: source=99997497 target=83459078". Switched to the -f file-mode pattern run_sql_file() already uses. -mem_limit=14gb was a c6a.2xlarge dev leftover that capped Impala well below available memory on larger hosts. Changed to -mem_limit=80%, Impala's native percentage syntax, per Cloudera's guidance.
|
Fixed. Two issues here. The bucket-UPSERT loop piped SQL into impala-shell over stdin without -f/-q, which always exits 0 on EOF regardless of query failure, so a failed batch logged as "exit=0" and the loader kept going past silently-incomplete buckets. Switched it to the same -f file-mode pattern the rest of the script already uses, so a real failure now aborts the load with a nonzero exit. Separately, Tested a deliberately broken UPSERT batch on both hosts and it now exits 1 instead of continuing silently. On EC2 c6a.8xlarge (~61.5 GiB), mem_limit resolves to 49.2 GB and Impala peaked around 17.5 GB loading the full 99,997,497-row dataset, which completed cleanly with an exact row count match. |
|
One more thing while I'm in this file: I'd like to do a follow-up fix. Will report back once it's verified. |
|
The run of Logs:
|
|
Results for Logs:
|
|
Results for Logs:
|
|
Results for Logs:
|
4.5.0 is hit by two CVEs fixed in 4.5.2: CVE-2026-56207, a critical SAML auth bypass, and CVE-2026-57866, an SSRF that leaks secrets. Reran the load and UPSERT-failure checks locally and against the full dataset on an EC2 c6a.8xlarge; row counts, exit codes and mem_limit resolution all match 4.5.0, so this is a drop-in bump.
|
Bumped to 4.5.2. Re-verified load and the failure-handling fix on EC2 with the full dataset. |
Adds an
impala-kudu/entry. Kudu holds the data; Impala runs the SQL. This is separate fromimpala/, which queries the source Parquet file directly.On
c6a.2xlargewith Ubuntu 24.04, the full 100M-row run loaded99,997,497rows into Kudu, matching the source. Load took3,800.123 s, and allocated persistent storage was47,958,401,024 bytes. All 43 queries x 3 tries have timings.The frozen validator passed 42 of 43 queries. Q4 was within the predeclared floating-point tolerance; exact checks also matched
COUNT(*),COUNT(UserID), andSUM(CAST(UserID AS DECIMAL(38,0))). Original Q41 and Q42 selected different rows at tied boundaries. Q41 still passed under the frozen rules; Q42 was the only failure. Tie-breaker diagnostics matched both. The timed SQL was unchanged.After
checkstarted verifying Kudu health, I reran the concurrent phase:0.100 QPS, error ratio0.016. The shared driver does not retain worker-level errors, so the failed concurrent query is unknown.Tagged
no-coldlike the existing Impala entry: the driver drops the host page cache but does not restart the stack between tries.data-sizereports allocated blocks for Kudu and Hive Metastore, excluding the source Parquet file.Closes #1090.