Skip to content

perf(telemetry): remove the typed Parquet writer ingestion regression #277

Description

@vishr

PR #272 review follow-up for parquet_writer.go: #272 (comment)

The three alternating Linux ARM64 ingestion-only repetitions against #270 reproduced a roughly 15% durable-throughput regression. Optimize the measured typed writer/copy/allocation path while retaining format 3, typed values, nanosecond UTC columns, complete trace ordering, durable acknowledgements, and conservative admission bounds.

Median baseline -> DuckDB 2: 484,312 -> 411,796 acknowledged rows/s; 3,826 -> 4,123 Go allocated bytes/row; 1.316 -> 1.478 client+server CPU cores; 36.54 -> 42.09ms export p95. CPU cores rose about 12%; CPU time per acknowledged row rose about 32%, since throughput also declined. All six runs had zero export errors and passed storage verification. The fixture is 16 shared-transport clients, 1,000 rows/export, equal signals, 4 CPUs/6GiB, ten-second intervals, in-process clients/server, no dashboard queries. It is not the harness mixed-p8 workload. ValueBytes was about 2% cumulative CPU in the first current profile; its per-key charge is accounting, not allocated bytes.

Validation: repeat baseline/current alternating runs at least three times with CPU/post-load heap profiles, then repeat realistic mixed-p8 and cooldown RSS measurements. Preserve deep typed/nested/binary round trips, shredded residual types, compaction and crash/corruption checks. Profiles are generated with FANOUT_BENCH_PROFILE_DIR and the transportbench build tag; include inspectable reports/artifacts with the final measurements.

Implementation is authorized in PR #272; this issue tracks the fix and its validation.

Activity

  1. vishr commented on Oct 1, 2026

    @vishr
    MemberAuthor

    Implemented reusable scalar buffers, per-column dispatch, and bounded shared-resource physical-cell reuse; unique resources still stream through the typed format-3 encoder. Added mixed shared/unique/null resource round trips across chunks and row groups, with allocation and compaction guards. Group admission is now 10ms with the same durable publication contract and bounded queues. Three alternating baseline/current profiling repetitions now measure median 493,050 to 586,813 durable rows/s (+19%), p95 35.68 to 31.75ms, zero export errors, and successful storage verification. The throughput regression is resolved on this fixture. CPU per row remains +26% (3.437 vs 2.722 CPU seconds/million rows), and allocations +5%; comparing occupied cores alone is not a per-row CPU comparison. The mixed-load run still fails saturated read capacity, and its adaptive sustained load is twice the old selection, so it is not presented as a matched capacity comparison.

    Implemented in ea13675; tracked in #277. Measurement report and reproduction. just check, the full Go race suite, and native Linux ARM64 full Go tests pass. Issue closure is linked to merge of #272.

  2. vishr commented on Oct 1, 2026

    @vishr
    MemberAuthor

    Still open after commit 3dd6f25. The historical +19% throughput fixture result is admission-limited; it does not establish that typed encoding CPU cost is resolved. The PR body and summary no longer close this issue.

    Three alternating baseline/current repetitions on native Linux ARM64, one CPU / 6 GiB, durable mode, sixteen in-process clients and 1,000 rows/export show median 414,700 → 321,598 acknowledged rows/s (-22.5%), 2.405 → 3.115 CPU seconds/million rows (+29.5%), and 3,746 → 3,943 Go bytes/row (+5.3%). All six have zero export errors and pass storage verification. CPU profiles cover the ten-second measurement interval; post-load heap profiles cover the full experiment. The cache follow-up does not change writer/admission code.

    report and reproduction. #278 separately tracks the remaining sustained dashboard-read capacity and peak RSS; cache stability or a smaller catalog should not be taken as a fix for this issue.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions