Skip to content

Add Apache Kudu - #2262

Open
KazukiKandaKK wants to merge 10 commits into
ClickHouse:mainfrom
KazukiKandaKK:add-apache-kudu
Open

KazukiKandaKK wants to merge 10 commits into
ClickHouse:mainfrom
KazukiKandaKK:add-apache-kudu

Conversation

@KazukiKandaKK

Copy link
Copy Markdown
Contributor

Adds an impala-kudu/ entry. Kudu holds the data; Impala runs the SQL. This is separate from impala/, which queries the source Parquet file directly.

On c6a.2xlarge with Ubuntu 24.04, the full 100M-row run loaded 99,997,497 rows into Kudu, matching the source. Load took 3,800.123 s, and allocated persistent storage was 47,958,401,024 bytes. All 43 queries x 3 tries have timings.

The frozen validator passed 42 of 43 queries. Q4 was within the predeclared floating-point tolerance; exact checks also matched COUNT(*), COUNT(UserID), and SUM(CAST(UserID AS DECIMAL(38,0))). Original Q41 and Q42 selected different rows at tied boundaries. Q41 still passed under the frozen rules; Q42 was the only failure. Tie-breaker diagnostics matched both. The timed SQL was unchanged.

After check started verifying Kudu health, I reran the concurrent phase: 0.100 QPS, error ratio 0.016. The shared driver does not retain worker-level errors, so the failed concurrent query is unknown.

Tagged no-cold like the existing Impala entry: the driver drops the host page cache but does not restart the stack between tries. data-size reports allocated blocks for Kudu and Hive Metastore, excluding the source Parquet file.

Closes #1090.

@alexey-milovidov alexey-milovidov added the machine:c6a.4xlarge PR benchmark machine: c6a.4xlarge (16 vCPU, 32 GB, AMD) — the default label Sep 27, 2026
@alexey-milovidov
alexey-milovidov deployed to benchmark-approval September 27, 2026 22:37 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

Results for impala-kudu are ready for: c6a.4xlarge.
The result files are committed as a6d958e.
Removed manually added result files: impala-kudu/results/20260927/c6a.2xlarge.json.

Logs:

@alexey-milovidov alexey-milovidov added machine:all PR benchmark on every machine type and removed machine:c6a.4xlarge PR benchmark machine: c6a.4xlarge (16 vCPU, 32 GB, AMD) — the default labels Sep 29, 2026
@alexey-milovidov alexey-milovidov self-assigned this Sep 29, 2026
@alexey-milovidov
alexey-milovidov deployed to benchmark-approval September 29, 2026 23:36 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

The run of impala-kudu on c6a.metal did not produce results.
The run of impala-kudu on c7a.metal-48xl did not produce results.
The run of impala-kudu on c8g.4xlarge did not produce results.
The run of impala-kudu on c8g.metal-48xl did not produce results.

Logs:

@github-actions

Copy link
Copy Markdown
Contributor

Results for impala-kudu are ready for: c6a.4xlarge.
The result files are committed as 5105250.

Logs:

@github-actions

Copy link
Copy Markdown
Contributor

Results for impala-kudu are ready for: c6a.2xlarge.
The result files are committed as e44ff35.

Logs:

@github-actions

Copy link
Copy Markdown
Contributor

Results for impala-kudu are ready for: c6a.xlarge.
The result files are committed as 7c859c1.

Logs:

@github-actions

Copy link
Copy Markdown
Contributor

The run of impala-kudu on c6a.large did not produce results.
The run of impala-kudu on t3a.small did not produce results.

Logs:

@alexey-milovidov

Copy link
Copy Markdown
Member

@KazukiKandaKK

Sometimes it fails to load the dataset:

row count mismatch: source=99997497 target=83459078

on c6a.metal.

The bucket-UPSERT loop piped SQL into impala-shell's stdin mode, which
exits 0 on EOF even when a query fails, masking OOM/syntax errors
behind "row count mismatch: source=99997497 target=83459078". Switched
to the -f file-mode pattern run_sql_file() already uses.

-mem_limit=14gb was a c6a.2xlarge dev leftover that capped Impala well
below available memory on larger hosts. Changed to -mem_limit=80%,
Impala's native percentage syntax, per Cloudera's guidance.
@KazukiKandaKK
KazukiKandaKK deployed to benchmark-approval October 1, 2026 11:59 — with GitHub Actions Active
@KazukiKandaKK

Copy link
Copy Markdown
Contributor Author

Fixed. Two issues here.

The bucket-UPSERT loop piped SQL into impala-shell over stdin without -f/-q, which always exits 0 on EOF regardless of query failure, so a failed batch logged as "exit=0" and the loader kept going past silently-incomplete buckets. Switched it to the same -f file-mode pattern the rest of the script already uses, so a real failure now aborts the load with a nonzero exit.

Separately, -mem_limit=14gb was a leftover from developing on the c6a.2xlarge dev machine, where even a single bulk UPSERT of the full dataset OOM'd at that limit (that's why the loader is bucketed in the first place). The value was never revisited, so on c6a.metal it capped Impala far below the 384 GiB available, and a real query went over it mid-load, which is what produced the mismatch. Changed it to -mem_limit=80%, Impala's native percentage syntax; Cloudera's FAQ recommends roughly that figure for dedicated Impala nodes.

Tested a deliberately broken UPSERT batch on both hosts and it now exits 1 instead of continuing silently. On EC2 c6a.8xlarge (~61.5 GiB), mem_limit resolves to 49.2 GB and Impala peaked around 17.5 GB loading the full 99,997,497-row dataset, which completed cleanly with an exact row count match.

@KazukiKandaKK

Copy link
Copy Markdown
Contributor Author

One more thing while I'm in this file: I'd like to do a follow-up fix. Will report back once it's verified.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

The run of impala-kudu on c6a.large did not produce results.
The run of impala-kudu on c8g.4xlarge did not produce results.
The run of impala-kudu on c8g.metal-48xl did not produce results.
The run of impala-kudu on t3a.small did not produce results.

Logs:

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Results for impala-kudu are ready for: c6a.metal, c7a.metal-48xl.
The result files are committed as 1a5d0b9.

Logs:

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Results for impala-kudu are ready for: c6a.2xlarge, c6a.4xlarge.
The result files are committed as cdc35cf.

Logs:

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Results for impala-kudu are ready for: c6a.xlarge.
The result files are committed as 72568e4.

Logs:

4.5.0 is hit by two CVEs fixed in 4.5.2: CVE-2026-56207, a critical
SAML auth bypass, and CVE-2026-57866, an SSRF that leaks secrets.
Reran the load and UPSERT-failure checks locally and against the full
dataset on an EC2 c6a.8xlarge; row counts, exit codes and mem_limit
resolution all match 4.5.0, so this is a drop-in bump.
@KazukiKandaKK

KazukiKandaKK commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Bumped to 4.5.2. Re-verified load and the failure-handling fix on EC2 with the full dataset.

This branch is waiting to be deployed

1 waiting deployment
benchmark-approval — 9b7c76c5 Waiting Oct 2, 2026 by KazukiKandaKK via launch #615
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

machine:all PR benchmark on every machine type

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add Apache Kudu

3 participants