Skip to content

Group fixed-route MoE tokens and reuse expert weight tiles - #78

Open
Happymic wants to merge 1 commit into
mainfrom
feat/fixed-route-expert-grouping
Open

Happymic wants to merge 1 commit into
mainfrom
feat/fixed-route-expert-grouping

Conversation

@Happymic

@Happymic Happymic commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator

What changed

  • Add an opt-in tile-major HBM matrix layout so each 64 x 64 weight tile is physically contiguous and the generated stride matches that layout.
  • Add compact expert-major VRAM gather and fixed-route grouped expert lowering. Tokens routed to the same known expert share one gate/up/down weight load instead of reloading the same weights for every routed pair.
  • Add an opt-in two-panel ping-pong lowering path using existing H_PREFETCH_M and matrix instructions, with strict Matrix-SRAM capacity checks.
  • Restore stage attribution to expert_weight_prefetch after dynamic expert-address calculation.
  • Add CI coverage for storage layout, compact grouping, one-load-per-expert behavior, stage attribution, invalid grouping, and ping-pong capacity/order.

Scope and compatibility

This is a compiler-only, opt-in change. It does not add or rename opcodes, change RTL, or modify the default row-major/per-pair/blocking path. A direct comparison against main produced the same 74 non-comment instructions and the same SHA256 for the default path.

Grouped lowering is intentionally limited to fixed-route or trace-replay programs where expert groups are known when code is generated. It does not claim runtime grouping for device-selected routes; that still needs a runtime sort/segment mechanism.

The rejected experimental deep-ring and cross-expert panel-pool implementations are not included.

Validation

  • Python 3.12 minimal CI environment: 10 passed in test_moe_expert_grouping.py.
  • MoE compiler guards: 36 passed.
  • Full aten/tests: 94 passed; 3 failures reproduce unchanged on origin/main (Transformers router API drift, an existing packed-router comment assertion, and the existing slow quantization threshold).
  • git diff --check: clean.
  • No developer-local absolute paths or experiment artifacts are included.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant