Skip to content

[Bug] Schema changes on one HStore Server can stay invisible on the others indefinitely #3235

Description

@bitflicker64

Bug Type (问题类型)

logic (逻辑设计问题)

Before submit

  • 我已经确认现有的 Issues 与 FAQ 中没有相同 / 重复问题 (I have confirmed and searched that there are no similar problems in the historical issue and documents)

Environment (环境信息)

Expected & Actual behavior (期望与实际表现)

With several Servers on HStore, a schema change made through one Server should become visible on the others within a bounded time. Today some changes never do.

CachedSchemaTransactionV2 keeps JVM-wide per-graph schema caches with expiry turned off. Remote Servers only drop them when a PD event arrives on HUGEGRAPH/{cluster}/EVENT/GRAPH/SCHEMA/CLEAR:

  1. updateSchema() publishes nothing, on purpose, because status flips from background jobs would cause a broadcast storm. That covers every structural update (property append, a new index on a base label) and every status flip (REBUILDING, INVALID, DELETING). Remote caches keep the old element.
  2. removeSchema() and store clear/truncate publish only when task.sync_deletion=false.
  3. An event that is published can still be lost. After fix(pd): keep KvClient watches alive after reconnect failures #3157 the PD watch reconnects, but a new subscription starts from "now", so nothing sent during a PD restart or leader change is replayed.

Measured on the setup above: warm every schema cache on all three Servers, wait 22 s, then append user_data to a property key through server-0.

Change on server-0 server-1 server-2
add property key (event published) visible after 0.05 s visible after 0.06 s
append user_data to it (no event) still old after 60 s still old after 60 s
append again right after a PD pod restart still old after 60 s still old after 60 s

A remote Server only recovers when something else clears its cache: a later add event, a restart, or the name-miss reload in getSchema(type, name) when the task scheduler looks up ~task after an earlier clear. With enough traffic the bug looks intermittent, which matches #2617.

Why the one-line fix is not enough

Making updateSchema() notify (the TODO) turns M status flips on N Servers into M x (N-1) full remote cache clears, each followed by a reload. An index rebuild flips status 2 to 3 times per label, so a large rebuild clears every cache thousands of times. The analysis and numbers are in #3205.

Proposed fix

Each graph gets one opaque version token in PD, at HUGEGRAPH/{cluster}/SCHEMA_VERSION/{graphSpace}/{graph}. CachedSchemaTransactionV2 writes a fresh token after every committed schema change: add, update, status flip, remove and clear. Each Server runs one small task per open graph that reads the token every schema.sync.reconcile_interval seconds (default 10). When it changed, the task clears that graph's schema caches and remembers the value it read before clearing. Existing events stay as they are and still give fast propagation for adds. A remote Server therefore sees any change within one interval plus one PD read after the token write, even if the watch is down or events were lost. A burst of M changes costs each Server at most one clear per interval, so cluster-wide clears are N x (ceil(D/T) + 1) for a burst lasting D seconds, whatever M is. A generation counter on the shared cache attachment stops a read that started before a clear from putting stale data back after it. schema.sync.enabled=false restores today's behaviour exactly.

Open questions from #3205, answered with defaults the maintainer can override

The questions in #3205 have had no maintainer answer yet. The PR uses these defaults. Each can be changed or reverted without touching the design.

# Question Default in the PR How to override
Q1 HBase multi-server mode Out of scope. HBase uses the V1 path with no PD and is marked for removal in 2.0 A counter-backed version store can be added later behind the same interface
Q2 Reserve a HugeType code for the version register Not needed; the PD key is a plain string Reserve one if Q1 is revisited
Q3 Reconcile interval 10 s per graph, schema.sync.reconcile_interval, range 0 to 3600; 0 turns polling off for that graph Set 5 for a 5 s bound at twice the poll cost, or any other value per graph
Q4 PD read and watch semantics Checked on master: KV get/put/scan redirect to the leader (KvServiceGrpcImpl), a missing key reads as "", the watch has no replay. The design needs only leader-routed point reads Nothing to change if PD semantics change; a stale read only delays one tick
Q5 Retire the legacy CLEAR event key Not started. Events are emitted exactly as today Separate change, together with the Store parser fix
Q6 Dead notifyGraphVertexCacheClear/notifyGraphEdgeCacheClear puts Left as they are Separate change behind a flag
Q7 Store-side SchemaDriver parser Left as it is; Store nodes see the same events as today Separate hugegraph-struct change and Store release
Q8 Delete the task.sync_deletion gate Kept. Removal now writes the token unconditionally, so the gate only decides whether an extra event goes out Delete it when the legacy events are retired
Q9 Replace the name-miss full reload with point repair Kept. This design has no per-id deltas, so name misses do not become routine About 15 lines mirroring V1, in a later change
Q10 SchemaStatus audit Partial: index queries are gated on status().ok(); a flip into REBUILDING or DELETING now reaches peers within one interval (today it never does) Lower the interval, or add a fast lane later
Q11 DDL error contract when the signal write fails A failed token write never fails the DDL call. It logs a WARN and is retried on each reconcile tick. The add path keeps today's behaviour for its event put. If the JVM dies before the retry lands, that one change stays unsignalled until the next schema change of that graph Make the write strict with a small change if callers prefer an error

Related: discussion #3205 (design and analysis), #2617 (original report), #3011 (the current event path), #3157 (watch recovery, no replay).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions