Skip to content

core: Hash each WASM module once instead of once per data source - #6732

Open
madumas wants to merge 1 commit into
graphprotocol:masterfrom
ellipfra:hash-module-once
Open

madumas wants to merge 1 commit into
graphprotocol:masterfrom
ellipfra:hash-module-once

Conversation

@madumas

@madumas madumas commented Oct 9, 2026

Copy link
Copy Markdown

Fixes #6731.

SubgraphInstance::new_host computed keccak256 over the full module bytes for
every data source before looking up module_cache. Data sources created from
the same template share one Arc of module bytes, so on a subgraph with
millions of dynamic data sources this rehashed the same bytes millions of times
at every runner start. On the subgraph from #6731 (Uniswap v3 on Base,
~1.77M data sources created from one template, 66 KB module), startup goes
from 238 s to 11 s (measurements below).

What changes

  • New core/src/subgraph/context/instance/module_cache.rs with a small
    ModuleCache<R> that replaces the HashMap<[u8; 32], Sender<T::Req>> field
    of SubgraphInstance:
    • senders: HashMap<[u8; 32], Sender<R>> — unchanged semantics: keccak256
      of the module contents → channel to the mapping thread;
    • hashes: HashMap<usize, (Arc<Vec<u8>>, [u8; 32])> — memo of the hash,
      keyed by Arc::as_ptr of the module bytes;
    • sender(&module_bytes, spawn) hashes (memoized), returns the cached sender
      or calls spawn and caches its sender.
  • new_host calls self.module_cache.sender(&module_bytes, |bytes| T::spawn_mapping(…))
    instead of hashing inline. Everything else in new_host is unchanged.

Why this is safe

  • The cache key is still the keccak256 of the module contents. The memo
    only avoids recomputing a hash; it never decides which module a data source
    gets. Equal bytes in a different Arc are hashed again and map to the same
    sender, as before.
  • No address reuse (ABA). The memo holds a clone of the Arc, so the
    allocation cannot be freed and its address cannot be reused for other bytes
    while the entry exists. A hit also checks Arc::ptr_eq against the stored
    Arc.
  • No mutation. Module bytes are Arc<Vec<u8>> and never mutated in
    place (Arc::make_mut would allocate a new buffer, hence a new address).
  • No concurrency change. ModuleCache is a plain field used through
    &mut self, like the map it replaces.
  • Same behaviour otherwise. Same bytes, arguments and order passed to
    spawn_mapping; a failed spawn caches nothing and can be retried; a data
    source without a runtime is still skipped; a rebuilt SubgraphInstance
    starts with an empty cache.
  • Bounded memory. One memo entry (pointer + 32-byte hash + an Arc clone)
    per distinct module Arc. Templates and manifest data sources share their
    mapping's Arc, so this is bounded by the number of mappings in the
    manifest; the retained modules are the ones the senders already keep alive.

Tests

Unit tests in module_cache.rs (no database needed):

  • two_templates_two_hashes — distinct modules get distinct, correct keccak256
    hashes and distinct senders;
  • sources_of_one_template_share_hash_and_sender — clones of one Arc are
    hashed once and share the sender;
  • many_sources_hash_once_per_template — 100,000 lookups over two templates
    compute exactly two hashes and spawn two modules;
  • equal_unshared_bytes_give_same_hash_and_module — equal bytes in another
    Arc are rehashed and reuse the existing module;
  • dropped_bytes_are_not_confused_with_new_ones — 64 short-lived Arcs all
    hash to their correct keccak256.

Before / after

Same subgraph (1,770,112 data sources), same machine, database and build
settings, one startup at a time, measured from the start of the runner
build to the first block processed. The two builds differ only by this change.

Segment Before (n = 2) After (n = 3)
load data sources from the store (until Data source count at start) 3.5 s 3.4–3.5 s
create one host per data source, build the runner 231.3–235.2 s 6.1–6.3 s
start the runner, build filters, first block (until Start processing block) 1.3 s 1.3 s
Total 236.1 / 240.0 s 10.7–11.0 s

Measured with this exact patch applied on a v0.45.0-based build; the code
path is unchanged on master.

After the change, startup on this subgraph is dominated by creating the hosts
(~6 s) and reading the data sources from the store (~3.5 s).

Checks run locally on this branch: cargo fmt --all -- --check, cargo clippy --all-targets with -D warnings (no warnings), cargo check --release, cargo test -p graph-core module_cache (5/5), the workspace unit tests (all pass except two gnd formatter/codegen tests that fail identically on master in my environment) and the runner tests (16/16).

`SubgraphInstance::new_host` computed keccak256 over the full module bytes
for every data source before looking up `module_cache`. Data sources
created from the same template share one `Arc` of module bytes, so on a
subgraph with millions of dynamic data sources this rehashed the same
bytes millions of times at every runner start.

Memoize the hash by the identity of the `Arc` holding the bytes, keeping a
clone of the `Arc` so its address cannot be reused while the entry exists.
The `module_cache` key is still the keccak256 of the module contents.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Subgraph startup rehashes the WASM module once per data source (minutes for subgraphs with millions of dynamic data sources)

1 participant