Skip to content

Add Openvino support - #1461

Open
Supreet Singh Palne (spalne) wants to merge 2 commits into
mainfrom
user/spalne/OpenVINO
Open

Supreet Singh Palne (spalne) wants to merge 2 commits into
mainfrom
user/spalne/OpenVINO

Conversation

@spalne

@spalne Supreet Singh Palne (spalne) commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Adds OpenVINO (Intel NPU) support for winml perf --runtime ort-genai: Qwen3 context/iterator stages run on the OpenVINO NPU, with a new --openvino-config option for custom NPU load settings (bundles are cached per config).
Fixes garbage output from precompiled OpenVINO models (weight-shared stages couldn't load their weights) and shows the generated response in the perf report and JSON.
Verified on Intel NPU with Qwen3-0.6B: correct answers, ~1.5 s load, ~45–53 tokens/sec .

winml-intel

@spalne
Supreet Singh Palne (spalne) marked this pull request as ready for review October 7, 2026 21:03
@spalne
Supreet Singh Palne (spalne) requested a review from a team as a code owner October 7, 2026 21:03
@staticmethod
def _shared_group(tmp_path: Path, ep: str) -> tuple[GenaiSession, list, dict]:
import numpy as np
import onnx
src_onnx = self._bundle_dir / onnx_filename
for location in get_external_data_files(src_onnx):
src, dst = src_onnx.parent / location, compiled_dir / location
if dst.exists():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Refresh copied weight sidecars when recompiling.

On Windows without symlink privileges, the fallback below copies the external weights into _compiled. After the source weights change, _process_shared_group can successfully recompile the stages, but this dst.exists() check leaves the previous weight copies in place. The new weightless EPContexts then load the old weights, which can produce incorrect output or fail to load if the layouts changed. Please refresh non-symlink copies when the source changes, and cover the copy-fallback/recompile case.

raise click.UsageError("--openvino-config requires --ep openvino.")
import hashlib

config_digest = hashlib.sha256(openvino_config.read_bytes()).hexdigest()[:12]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Include resolved relative paths in the OpenVINO bundle cache key.

build_npu_load_config resolves a relative WEIGHTS_PATH against the config file's parent directory, but this digest hashes only the file bytes. Two identical configs in different directories can therefore select different weights directories while sharing the same bundle cache key. The second run takes the cache-hit return below and silently keeps the first config's resolved weights path. Please key the cache on the effective configuration (including resolved paths), or include the config's resolved location in its identity.

Comment thread npu_config_npuw.json
@@ -0,0 +1,8 @@
{

@DingmaomaoBJTU Qiong Wu (qiowu) (DingmaomaoBJTU) Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P0] Reuse the existing EP provider-options path instead of adding a one-off flag.

winml perf already exposes repeatable --ep-options KEY=VALUE, and the Qwen3 GenAI stage options for QNN, VitisAI, and OpenVINO are built together in qwen3/genai.py. Could this follow that pattern rather than introduce a separate --openvino-config flag and root-level JSON file? --ep-options does not parse a standalone JSON object, but it can carry JSON as the value of OpenVINO's load_config, for example:

--ep-options 'load_config={"NPU":{"NPU_TURBO":"YES","NPU_QDQ_OPTIMIZATION":"YES","NPU_COMPILER_TYPE":"PLUGIN","CACHE_MODE":"OPTIMIZE_SPEED"}}'

The current ort-genai command path does not appear to forward ep_options into GenaiSession, so please wire the existing option through to the context/iterator stage provider options and wrap the NPU properties in openvino.load_config there. Keep the OpenVINO-specific mapping/default alongside the existing provider-option builders; explicit options should override defaults. If this configuration is later intended for all OpenVINO NPU model paths, the shared provider/session layer would be a better home.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants