Repository navigation
Add Openvino support - #1461
Add Openvino support#1461Supreet Singh Palne (spalne) wants to merge 2 commits into
Conversation
| @staticmethod | ||
| def _shared_group(tmp_path: Path, ep: str) -> tuple[GenaiSession, list, dict]: | ||
| import numpy as np | ||
| import onnx |
| src_onnx = self._bundle_dir / onnx_filename | ||
| for location in get_external_data_files(src_onnx): | ||
| src, dst = src_onnx.parent / location, compiled_dir / location | ||
| if dst.exists(): |
There was a problem hiding this comment.
[P1] Refresh copied weight sidecars when recompiling.
On Windows without symlink privileges, the fallback below copies the external weights into _compiled. After the source weights change, _process_shared_group can successfully recompile the stages, but this dst.exists() check leaves the previous weight copies in place. The new weightless EPContexts then load the old weights, which can produce incorrect output or fail to load if the layouts changed. Please refresh non-symlink copies when the source changes, and cover the copy-fallback/recompile case.
| raise click.UsageError("--openvino-config requires --ep openvino.") | ||
| import hashlib | ||
|
|
||
| config_digest = hashlib.sha256(openvino_config.read_bytes()).hexdigest()[:12] |
There was a problem hiding this comment.
[P2] Include resolved relative paths in the OpenVINO bundle cache key.
build_npu_load_config resolves a relative WEIGHTS_PATH against the config file's parent directory, but this digest hashes only the file bytes. Two identical configs in different directories can therefore select different weights directories while sharing the same bundle cache key. The second run takes the cache-hit return below and silently keeps the first config's resolved weights path. Please key the cache on the effective configuration (including resolved paths), or include the config's resolved location in its identity.
| @@ -0,0 +1,8 @@ | |||
| { | |||
There was a problem hiding this comment.
[P0] Reuse the existing EP provider-options path instead of adding a one-off flag.
winml perf already exposes repeatable --ep-options KEY=VALUE, and the Qwen3 GenAI stage options for QNN, VitisAI, and OpenVINO are built together in qwen3/genai.py. Could this follow that pattern rather than introduce a separate --openvino-config flag and root-level JSON file? --ep-options does not parse a standalone JSON object, but it can carry JSON as the value of OpenVINO's load_config, for example:
--ep-options 'load_config={"NPU":{"NPU_TURBO":"YES","NPU_QDQ_OPTIMIZATION":"YES","NPU_COMPILER_TYPE":"PLUGIN","CACHE_MODE":"OPTIMIZE_SPEED"}}'
The current ort-genai command path does not appear to forward ep_options into GenaiSession, so please wire the existing option through to the context/iterator stage provider options and wrap the NPU properties in openvino.load_config there. Keep the OpenVINO-specific mapping/default alongside the existing provider-option builders; explicit options should override defaults. If this configuration is later intended for all OpenVINO NPU model paths, the shared provider/session layer would be a better home.
Adds OpenVINO (Intel NPU) support for
winml perf --runtime ort-genai: Qwen3 context/iterator stages run on the OpenVINO NPU, with a new--openvino-configoption for custom NPU load settings (bundles are cached per config).Fixes garbage output from precompiled OpenVINO models (weight-shared stages couldn't load their weights) and shows the generated response in the perf report and JSON.
Verified on Intel NPU with Qwen3-0.6B: correct answers, ~1.5 s load, ~45–53 tokens/sec .