Skip to content

RESOURCE_LOCKED is treated as an ordinary load failure, so the destination retries instead of backing off #64

Description

@anoop-narang

What's missing

When another operation holds the write lock on a managed table, the Hotdata API responds:

  • 409, with the error code RESOURCE_LOCKED in the body
  • a Retry-After header (currently 5)
  • a message naming the resource that is busy

The code is distinct from a plain CONFLICT precisely so a client can tell transient contention (back off and retry) from a permanent duplicate-resource conflict (don't). This destination doesn't act on either signal — a locked load is just a failed load job, so what happens next is dlt's generic load-job retry: it retries the job, holding its data, at its own cadence rather than the one the server asked for.

RESOURCE_LOCKED does not appear anywhere in this repo.

Precedent in the ecosystem

hotdata-materialized already handles this — hotdata_materialized/store.py catches the conflict and waits the advertised interval, retrying a bounded number of times, with the comment that a data load can still hold the table lock when it runs.

That precedent is a single administrative call, though, not a fan-out of parallel load jobs, so it may not transfer directly — see the design question below.

Where it bites

A large backfill into one managed table with LOAD__WORKERS greater than 1. Writes to a single table are serialized, so only one load job can be admitted at a time and the rest come back RESOURCE_LOCKED. With eight workers the destination spends most of a large load being refused, retrying on dlt's schedule rather than the server's, and each queued job holds its data while it waits.

Observed downstream on a pipeline re-reading a few thousand objects (~1 GB) into a single table: sustained lock rejections for the duration of the load, and the job recycled several times without completing.

I'd be careful attributing memory pressure to this specifically — with parquet files capped at 50 MB, eight in-flight jobs only accounts for a few hundred MB, and the observed peak was several times that, so some of it is likely elsewhere (normalize, or allocator high-water). What isn't in question is that the client generates writes without reference to a back-off signal it is already being sent.

The design question

Two shapes, and I don't have a strong view:

  1. Honour Retry-After inside the load job — wait the advertised interval and retry, so contention resolves without surfacing. Simple, matches the precedent above. But N parallel jobs each sleeping independently still wake into the same lock, so it may need jitter or a shared gate to avoid a thundering herd.
  2. Surface a typed exception so the caller can throttle the producer instead of retrying the consumer. Pushes the decision to whoever knows the workload shape, at the cost of every caller implementing it.

Either way the invariant worth having: a lock 409 is not a permanent failure, and should not be retried faster than the server has said is useful.

Not covered here

Whether a caller should also cap its own write concurrency when loading into a single table. That's the caller's side and is tracked separately.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions