Skip to content

[Bug] A Server that starts while PD is briefly unreachable serves REST but loses its Gremlin binding for the life of the process #3228

Description

@bitflicker64

Bug Type (问题类型)

gremlin (结果不合预期)

Before submit

  • I have confirmed and searched that there are no similar problems in the historical issue and documents

Environment (环境信息)

Expected & Actual behavior (期望与实际表现)

When a Server starts during a PD rolling restart (a normal event under Kubernetes: rotating the PD REST secret rolls PD and Server together), the expected behavior is that the graph becomes fully usable once PD is reachable again.

Measured on 2026-09-22, 4 of 12 Server starts that overlapped a PD roll (0 of 3 in a Server-only roll):

// startup, PD momentarily unreachable
o.a.h.p.c.AbstractClient - connect to hugegraph-pd-0...:8686,... with error
o.a.t.g.s.u.DefaultGraphManager - Graph [DEFAULT-hugegraph] configured at
  [/hugegraph-server/conf/graphs/hugegraph.properties] could not be instantiated
  and will not be available in Gremlin Server.
// 8 s later the REST layer opens the same graph fine
o.a.h.StandardHugeGraph - Init system info for graph 'DEFAULT-hugegraph'

Afterwards the process passes readiness (/versions), REST reads and writes work, but every POST /gremlin fails with Could not rebind [graph] to [DEFAULT-hugegraph] as [DEFAULT-hugegraph] not in the Graph or TraversalSource global bindings, permanently. It persisted for the life of the pod; deleting the pod fixed it (2 of 2).

Cause as read from source: Gremlin Server gets graphs either from the static settings at construction (HugeGremlinServer.prepare) or through the GRAPH_CREATE event handled by ContextGremlinServer.injectGraph. Only the create paths fire that event (GraphManager.java:1460 and 1620, both in create methods); the startup load of existing graphs never does. So a graph whose static instantiation failed once is never offered to Gremlin again, even though the REST GraphManager holds it.

This is a different trigger from #3137/#3138 (create-then-query divergence across replicas) and from #3151 (missed PD watch events): here the graph is in the local static config and is present in the REST layer of the same process; only the embedded Gremlin Server missed it. A reconciliation as proposed in #3151 would presumably cover it if it also re-checks Gremlin bindings for locally-open graphs; a smaller fix is to fire the same injection after a successful startup load when Gremlin does not hold the graph, or to retry the failed static instantiation.

// Query URL
POST http://:8080/gremlin
{"gremlin":"graph.traversal().V().limit(1).count()","aliases":{"graph":"DEFAULT-hugegraph"}}
// -> {"exception":"...","message":"Could not rebind [graph] to [DEFAULT-hugegraph] ..."}

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions