Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
117 changes: 117 additions & 0 deletions kylin/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,117 @@
# Apache Kylin 5.0.2-GA

Standalone all-in-one Docker image (`apachekylin/apache-kylin-standalone:5.0.2-GA`),
bundling HDFS, YARN, Zookeeper, Hive, Spark `3.3.0-kylin-5.2.2`, and Gluten
`1.3.0-kylin-250110` (a native ClickHouse-based execution engine Kylin
layers on top of Spark). No cube is built. Every query below is a raw
pushdown query against the loaded `CLICKBENCH.HITS` table, so this does
not measure Kylin's core value proposition (pre-aggregated cubes).

Kylin 4.x is marked `retired` upstream; `5.0.2-GA` is the Docker Hub
image tag marked "Recommended for users" (no `5.0.3`/`5.0.4` standalone
image exists). The container is not restarted between query tries
(`BENCH_RESTARTABLE=no`), so this is tagged `no-cold`. Load time was
`421.406s`.

## Memory: `-m 10G`, the vendor default

Docker Hub's `5.0.2-GA` example specifies `-m 10G`; the quickstart
guide's `-m 8G` example is for the older `5.0-beta` tag, not this one.
Left at the vendor default. Also tested `install`/`start` only (no data
load) on `t3a.small` (2 GiB, fails its own readiness poll from swap
pressure, not an OOM kill) and `c6a.large` (4 GiB, succeeds).
`c6a.xlarge`, `c6a.4xlarge`, `c6a.metal`, and `c7a.metal-48xl` are
untested.

## Canary config: two keys set, likely no-ops per source

`./start` sets `kylin.canary.sparder-context-canary-enabled=false` and
`kylin.canary.sqlcontext-enabled=false`, then restarts Kylin once during
initial setup. In 5.0.2's source, `sparder-context-canary-enabled`
matches no config key anywhere; `sqlcontext-enabled` is read by
`KapConfig.getSparkCanaryEnable()`, which already defaults to `false`,
and gates whether `SparkContextCanary` starts (`SparderConfiguration.init()`),
unless an earlier `spark.local=true` check short-circuits that method
first, in which case the gate is never reached at all. Which path
actually runs in this container wasn't checked at runtime, so setting
these two keys to `false` is likely redundant rather than confirmed
inert; they're kept rather than removed.

## No Parquet rewrite needed

This build's Spark reads the official `hits.parquet`'s `EventDate`
column (Parquet `INT32`/`UINT_16`) without any rewrite, confirmed by a
separate `SELECT COUNT(*), MIN(EventDate), MAX(EventDate)` returning
`99,997,497` rows and the correct `2013-07-02`..`2013-07-31` range.
`./load` copies the file into HDFS unchanged, so `Data size` equals the
source file's byte length.

## Query result cache

Lookup stays on (`kylin.query.cache-enabled` defaults `true`, and this
`query` script doesn't force pushdown). The pushdown success-store path
defaults off (`kylin.query.pushdown.cache-enabled=false`), so no
successful pushdown result is ever cached here. The `ResourceLimitExceededException`
failure-cache path (which bypasses `kylin.query.exception-cache-enabled`)
is never thrown anywhere in 5.0.2's non-test source, so it doesn't
apply to this build's Gluten OOM failures; the ordinary exception cache
(`kylin.query.exception-cache-enabled`) also defaults off. Spark/Gluten's
own internal caching wasn't examined.

## Load row-count check

A post-run `COUNT(*) FROM clickbench.hits` returned `99997497` (exact
match), checked after the benchmark, not during it.

## Known failures (cube-less pushdown, Kylin 5.0.2-GA)

38 of 43 queries completed on all 3 tries. 5 recorded as `null`:

- **Q9, Q10**: Gluten aborts with `Memory limit exceeded ... maximum: 1.00 GiB`
on a grouped `COUNT(DISTINCT UserID)` with no `WHERE` filter. Other
`COUNT(DISTINCT)` queries succeed (Q5 unfiltered/ungrouped;
Q11/Q12/Q14/Q23 filtered before grouping); the distinguishing factor
wasn't isolated.
- **Q19, Q29, Q40**: Calcite rejects a `SELECT`-list alias (`m`/`k`/`Src`,
none submitted uppercase or quoted) referenced in `GROUP BY`
(`Column 'M'/'K'/'SRC' not found`), consistent with case-folding
during resolution. This is a validation-time failure before Gluten or Spark
runs. Q19's submitted `extract(minute FROM EventTime) AS m` is
rewritten by Kylin to `MINUTE(EventTime) AS m` before the alias fails,
so the error SQL differs from the submitted SQL.

`query` treats non-200 HTTP and in-band `isException:true` (HTTP 200)
as failures; `isPartial` isn't inspected.

## OFFSET

Q39/41/42/43 contain `OFFSET` and complete; Q40 also contains it but
fails on the alias issue first. Spark itself only gained `OFFSET`
support in 3.4.0 ([SPARK-28330](https://issues.apache.org/jira/browse/SPARK-28330));
this build's Spark is `3.3.0-kylin-5.2.2`, and how Kylin accepts it
despite that wasn't determined. Completion doesn't confirm the skipped
rows were correct (no independent check against a known-correct offset).

## Concurrent-QPS

10 connections, 600s: `0.058` QPS / `0.167` error ratio. Per
`bench_concurrent_qps` (`qps=ok/600`, `error_ratio=err/(ok+err)`,
`%.3f`-rounded), 35/7 is the only pair matching both. All 7 traced by
traceId: 3 match the sequential failures above; 4 are queries that
passed sequentially but failed only under concurrency with
`QueryInterruptChecker: ... Interrupted at the stage of collecting
result`; trigger not isolated.

## amd64 only

`5.0.2-GA`'s Docker Hub manifest lists amd64 only; `./install` fails
outright on arm64 hosts (`c8g.4xlarge`, `c8g.metal-48xl`).

## `./start` idempotency caveat

`./start` skips the canary-key block whenever the container already
exists, assuming prior setup completed. If an earlier readiness
timeout left the container existing-but-unconfigured, a later `./start`
would skip the block without having applied it. Whether this run hit
that path isn't verifiable from the committed files (the harness
discards `./start`'s stdout).
12 changes: 12 additions & 0 deletions kylin/benchmark.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
#!/bin/bash
export BENCH_DOWNLOAD_SCRIPT="download-hits-parquet-single"

# Full container restart between tries is too slow (bundles SSH, MySQL,
# Zookeeper, Hadoop, Hive, Kylin, Spark, Gluten). Not a durability
# issue: BENCH_DURABLE stays at its default "yes".
export BENCH_RESTARTABLE=no

# First boot of the full stack; 900s is generous headroom.
export BENCH_CHECK_TIMEOUT=900

exec ../lib/benchmark-common.sh
8 changes: 8 additions & 0 deletions kylin/check
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
#!/bin/bash
# Same readiness probe ./start uses.
KYLIN_TOKEN=$(printf 'ADMIN:KYLIN' | base64)
REQUEST_AUTH_HEADER_NAME=Authorization
HTTP_AUTH_SCHEME=Basic
curl -sf -m 5 http://localhost:7070/kylin/api/user/authentication \
-X POST -H "${REQUEST_AUTH_HEADER_NAME}: ${HTTP_AUTH_SCHEME} ${KYLIN_TOKEN}" \
-H 'Content-Type: application/json' >/dev/null
144 changes: 144 additions & 0 deletions kylin/create.sql
Original file line number Diff line number Diff line change
@@ -0,0 +1,144 @@
CREATE DATABASE IF NOT EXISTS clickbench;
USE clickbench;

DROP VIEW IF EXISTS hits;
DROP TABLE IF EXISTS hits_raw;

CREATE EXTERNAL TABLE hits_raw (
WatchID bigint,
JavaEnable smallint,
Title string,
GoodEvent smallint,
EventTime bigint,
EventDate int,
CounterID int,
ClientIP int,
RegionID int,
UserID bigint,
CounterClass smallint,
OS smallint,
UserAgent smallint,
URL string,
Referer string,
IsRefresh smallint,
RefererCategoryID smallint,
RefererRegionID int,
URLCategoryID smallint,
URLRegionID int,
ResolutionWidth smallint,
ResolutionHeight smallint,
ResolutionDepth smallint,
FlashMajor smallint,
FlashMinor smallint,
FlashMinor2 string,
NetMajor smallint,
NetMinor smallint,
UserAgentMajor smallint,
UserAgentMinor string,
CookieEnable smallint,
JavascriptEnable smallint,
IsMobile smallint,
MobilePhone smallint,
MobilePhoneModel string,
Params string,
IPNetworkID int,
TraficSourceID smallint,
SearchEngineID smallint,
SearchPhrase string,
AdvEngineID smallint,
IsArtifical smallint,
WindowClientWidth smallint,
WindowClientHeight smallint,
ClientTimeZone smallint,
ClientEventTime bigint,
SilverlightVersion1 smallint,
SilverlightVersion2 smallint,
SilverlightVersion3 int,
SilverlightVersion4 smallint,
PageCharset string,
CodeVersion int,
IsLink smallint,
IsDownload smallint,
IsNotBounce smallint,
FUniqID bigint,
OriginalURL string,
HID int,
IsOldCounter smallint,
IsEvent smallint,
IsParameter smallint,
DontCountHits smallint,
WithHash smallint,
HitColor string,
LocalEventTime bigint,
Age smallint,
Sex smallint,
Income smallint,
Interests smallint,
Robotness smallint,
RemoteIP int,
WindowName int,
OpenerName int,
HistoryLength smallint,
BrowserLanguage string,
BrowserCountry string,
SocialNetwork string,
SocialAction string,
HTTPError smallint,
SendTiming int,
DNSTiming int,
ConnectTiming int,
ResponseStartTiming int,
ResponseEndTiming int,
FetchTiming int,
SocialSourceNetworkID smallint,
SocialSourcePage string,
ParamPrice bigint,
ParamOrderID string,
ParamCurrency string,
ParamCurrencyID smallint,
OpenstatServiceName string,
OpenstatCampaignID string,
OpenstatAdID string,
OpenstatSourceID string,
UTMSource string,
UTMMedium string,
UTMCampaign string,
UTMContent string,
UTMTerm string,
FromTag string,
HasGCLID smallint,
RefererHash bigint,
URLHash bigint,
CLID int
)
STORED AS PARQUET
LOCATION 'hdfs:///clickbench/hits_dir';

CREATE VIEW hits AS
SELECT
WatchID, JavaEnable, Title, GoodEvent,
CAST(from_unixtime(EventTime) AS TIMESTAMP) AS EventTime,
date_add(DATE '1970-01-01', EventDate) AS EventDate,
CounterID, ClientIP, RegionID, UserID, CounterClass, OS, UserAgent,
URL, Referer, IsRefresh, RefererCategoryID, RefererRegionID,
URLCategoryID, URLRegionID, ResolutionWidth, ResolutionHeight,
ResolutionDepth, FlashMajor, FlashMinor, FlashMinor2, NetMajor, NetMinor,
UserAgentMajor, UserAgentMinor, CookieEnable, JavascriptEnable,
IsMobile, MobilePhone, MobilePhoneModel, Params, IPNetworkID,
TraficSourceID, SearchEngineID, SearchPhrase, AdvEngineID, IsArtifical,
WindowClientWidth, WindowClientHeight, ClientTimeZone,
CAST(from_unixtime(ClientEventTime) AS TIMESTAMP) AS ClientEventTime,
SilverlightVersion1, SilverlightVersion2, SilverlightVersion3,
SilverlightVersion4, PageCharset, CodeVersion, IsLink, IsDownload,
IsNotBounce, FUniqID, OriginalURL, HID, IsOldCounter, IsEvent,
IsParameter, DontCountHits, WithHash, HitColor,
CAST(from_unixtime(LocalEventTime) AS TIMESTAMP) AS LocalEventTime,
Age, Sex, Income, Interests, Robotness, RemoteIP, WindowName,
OpenerName, HistoryLength, BrowserLanguage, BrowserCountry,
SocialNetwork, SocialAction, HTTPError, SendTiming, DNSTiming,
ConnectTiming, ResponseStartTiming, ResponseEndTiming, FetchTiming,
SocialSourceNetworkID, SocialSourcePage, ParamPrice, ParamOrderID,
ParamCurrency, ParamCurrencyID, OpenstatServiceName, OpenstatCampaignID,
OpenstatAdID, OpenstatSourceID, UTMSource, UTMMedium, UTMCampaign,
UTMContent, UTMTerm, FromTag, HasGCLID, RefererHash, URLHash, CLID
FROM hits_raw;
6 changes: 6 additions & 0 deletions kylin/data-size
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
#!/bin/bash
set -euo pipefail
# HDFS logical length of the loaded file (hdfs dfs -du -s column 1,
# not disk_space_consumed_with_all_replicas), not the host download.
sudo docker exec kylin-clickbench hdfs dfs -du -s /clickbench/hits_dir/hits.parquet \
| awk '{print $1}'
23 changes: 23 additions & 0 deletions kylin/install
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
#!/bin/bash
set -e

# Apache Kylin 5.0.2-GA standalone (bundles Hadoop 3.2.4/HDFS/YARN/
# Zookeeper/Hive 3.1.3/Spark 3.3.0-kylin-5.2.2 inside one container).
# Docker is required; no external Hadoop cluster.
# Version choice rationale: see README.md.
if ! command -v docker >/dev/null 2>&1; then
sudo apt-get update -y
sudo apt-get install -y docker.io
fi
sudo apt-get install -y curl jq

# Add a 6GB swapfile if none exists: the container is -m 10G on a
# 16GiB host, and Hive's CREATE TABLE/VIEW in ./load can push RAM.
if [ "$(swapon --show | wc -l)" -eq 0 ] && [ ! -f /swapfile ]; then
sudo fallocate -l 6G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
fi

sudo docker pull apachekylin/apache-kylin-standalone:5.0.2-GA
53 changes: 53 additions & 0 deletions kylin/load
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
#!/bin/bash
set -e

CONTAINER=kylin-clickbench

# --- Dataset: expects hits.parquet already downloaded, no column rewrite needed ---
# Official hits.parquet, unmodified. See README.md ("No Parquet
# rewrite needed").
if [ ! -f hits.parquet ]; then
echo "hits.parquet not found in cwd; run ../lib/download-hits-parquet-single first" >&2
exit 1
fi

# --- Stage the file into HDFS inside the container -----------------------
sudo docker cp hits.parquet "$CONTAINER":/tmp/hits.parquet
# Drop the host copy so docker cp and hdfs put do not occupy the disk twice.
rm -f hits.parquet
sync
sudo docker exec "$CONTAINER" bash -c '
set -e
hdfs dfs -mkdir -p /clickbench/hits_dir
hdfs dfs -put -f /tmp/hits.parquet /clickbench/hits_dir/hits.parquet
rm -f /tmp/hits.parquet
'

# --- Register the Hive external table + ClickBench-typed view ------------
sudo docker cp create.sql "$CONTAINER":/tmp/create.sql
sudo docker exec "$CONTAINER" hive -f /tmp/create.sql

# --- Load the Hive table's metadata into Kylin via REST API --------------
# Stock standalone credentials, encoded at runtime.
KYLIN_USER=ADMIN
KYLIN_PASS=KYLIN
AUTH_TOKEN=$(printf '%s:%s' "$KYLIN_USER" "$KYLIN_PASS" | base64)
HTTP_AUTH_SCHEME=Basic
REQUEST_AUTH_HEADER_NAME=Authorization
AUTH_HEADER="${REQUEST_AUTH_HEADER_NAME}: ${HTTP_AUTH_SCHEME} ${AUTH_TOKEN}"

# Create project; ignore already-exists.
curl -sf -X POST http://localhost:7070/kylin/api/projects \
-H "$AUTH_HEADER" -H 'Content-Type: application/json' \
-d '{"name":"clickbench","description":"ClickBench hits benchmark"}' \
>/dev/null 2>&1 || true

# Load table metadata (project/tables in the body).
load_resp=$(curl -sf -X POST \
"http://localhost:7070/kylin/api/tables" \
-H "$AUTH_HEADER" -H 'Content-Type: application/json' \
-d '{"project":"clickbench","tables":["CLICKBENCH.HITS"],"data_source_type":9}')
echo "table load response: $load_resp"

rm -f hits.parquet
sync
Loading
Loading