winml perf¶
Benchmark an ONNX model's latency and throughput on a target device.
When to use this¶
Use winml perf when you want a quantitative latency and throughput baseline for a model on a specific device, or when you need to compare the performance impact of different precision settings, execution providers, or batch sizes.
Synopsis¶
Flags¶
| Flag | Short | Type | Default | Description |
|---|---|---|---|---|
--model |
-m |
TEXT |
— | HuggingFace model ID or path to a local .onnx file. Required. With --runtime ort-genai, also accepts a prebuilt genai bundle directory, or a HuggingFace model ID that is auto-built into a bundle on demand. |
--runtime |
winml-ort\|ort-genai |
winml-ort |
Inference runtime. winml-ort benchmarks single-shot ONNX inference; ort-genai benchmarks an onnxruntime-genai bundle (LLM generation: time-to-first-token + decode tokens/sec). With ort-genai, a model ID that is not a bundle directory is auto-built into one before benchmarking. An explicit --ep or --device selects both the transformer build and runtime target; without an override, the auto-build defaults to QNN/NPU. Bundles are cached under ~/.cache/winml/, separately for each explicit EP/device target. GenAI cache controls are tracked in issue #1275. |
|
--task |
TEXT |
auto-detected | Explicit task override (e.g., image-classification). Inferred from the model if omitted. |
|
--iterations |
INTEGER |
100 |
Number of timed inference iterations used to compute statistics. | |
--warmup |
INTEGER |
10 |
Number of warm-up iterations run before timing begins; excluded from statistics. | |
--device |
-d |
auto\|cpu\|gpu\|npu |
auto |
Device to run the benchmark on. auto selects the highest-priority available device. |
--device-luid |
TEXT |
— | Pin a physical adapter within the resolved EP/device pair using its LUID from winml sys (0xHHHHHHHH_0xLLLLLLLL, case-insensitive). Requires the EP to expose that adapter's LUID. Not supported with --runtime ort-genai. |
|
--precision |
TEXT |
auto |
Precision mode applied during model build: auto, fp32, fp16, int8, int16, or compound forms such as w8a16. |
|
--ep |
TEXT |
— | Force a specific execution provider (e.g., qnn, dml, vitisai, openvino, cpu). Overrides the device-to-provider mapping. |
|
--ep-options |
KEY=VALUE (multiple) |
— | Runtime EP provider option forwarded to the inference session (e.g., --ep-options htp_performance_mode=burst). Repeatable. Applies to both HuggingFace model IDs and ONNX file inputs. When detail op-tracing automatically compiles a raw ONNX model, these options are also applied to that compilation. |
|
--output |
-o |
PATH |
~/.cache/winml/perf/<slug>/<timestamp>.json |
Output JSON file path for the benchmark report. |
--batch-size |
INTEGER |
1 |
Batch size used when generating synthetic input tensors. Ignored when --input-data is set. |
|
--input-data |
PATH |
— | Path to a .npz file of real input tensors to benchmark with instead of randomly generated inputs. The archive's keys must match the model's inputs exactly; dtypes are cast to the model's expected dtype (with a warning) to mirror normal inference. Not supported with --module, --runtime ort-genai, or composite (dual-encoder) models. |
|
--shape-config |
PATH |
— | Path to a JSON file containing shape overrides (e.g., {"height": 480, "width": 480}). Used for Hugging Face export and random input generation; ignored in --module mode and when --input-data is set. |
|
--input-specs |
PATH |
— | JSON input tensor specs to merge into the Hugging Face export config before benchmarking. Symbolic string dimensions infer dynamic axes. Ignored for pre-exported ONNX files and in --module mode. |
|
--export-config |
PATH |
— | JSON ONNX export config overrides to apply when perf builds a Hugging Face model before benchmarking. Ignored for pre-exported ONNX files and in --module mode. |
|
--dynamic-axes |
PATH |
— | JSON dynamic axes mapping for Hugging Face ONNX export, for example {"input_ids": {"0": "batch", "1": "sequence"}}. Ignored for pre-exported ONNX files and in --module mode. |
|
--quantize/--no-quantize |
flag | true |
Run quantization during model build (use --no-quantize to skip it). Useful for measuring the fp32 baseline. |
|
--use-cache/--no-use-cache |
flag | true |
Reuse persistent model build artifacts. --no-use-cache performs a fresh build in a temporary folder and discards it after benchmarking. |
|
--rebuild/--no-rebuild |
flag | false |
Force model rebuild even if a cached artifact already exists. | |
--module |
TEXT |
— | PyTorch module class name for per-module benchmarking (e.g., BertAttention). Builds and times each matching instance separately. See Load and export. |
|
--monitor/--no-monitor |
flag | false |
Show a live NPU/CPU utilization chart while the benchmark runs and include hardware metrics in the JSON report. With --runtime ort-genai, the monitor wraps the genai load + generation benchmark. |
|
--op-tracing |
basic\|detail |
— | Enable operator-level profiling. QNN detail tracing requires an EPContext model; a raw ONNX input is detected and compiled automatically with the required profiling options. | |
--compile / --no-compile |
flag | false |
Compile the model to EPContext binaries during build. QNN detail op-tracing enables this automatically for a raw ONNX input unless --no-compile or --skip-build was explicitly specified. For --runtime ort-genai on the NPU, --compile pre-compiles each QNN stage (in an isolated subprocess) before generation. |
|
--compile-timeout |
INTEGER |
300 |
(ort-genai) Max seconds to compile each EPContext stage before falling back to the original ONNX. Requires --compile. |
|
--prompt |
TEXT |
Explain the theory of relativity in simple terms. |
(ort-genai) Prompt text to generate from. Wrapped in the bundle's chat template unless --no-apply-template. |
|
--apply-template/--no-apply-template |
flag | true |
(ort-genai) Wrap --prompt in the bundle's chat template before timing. |
|
--max-new-tokens |
INTEGER |
128 |
(ort-genai) Number of new tokens to generate per iteration. |
How it works¶
winml perf loads the model through WinMLAutoModel — accepting both HuggingFace IDs and local ONNX files — then generates random input tensors from the model's I/O configuration. It runs the specified number of warm-up iterations (excluded from statistics) followed by the timed iterations, collecting per-sample latency. The final report includes mean, min, max, P50, P90, P95, P99, standard deviation, and throughput in samples per second. When --monitor is active, a hardware polling loop runs in parallel and records NPU / GPU utilization, CPU usage, and device memory alongside the timing data.
Both runtime reports include schema_version: 2 and a benchmark_info.runtime discriminator (winml-ort or ort-genai). Shared metadata such as model_id, running_model_path, device, ep, iterations, warmup, and timestamp uses the same field names where the concepts overlap; GenAI also keeps bundle_dir because the runnable artifact is a bundle directory.
When --memory is enabled, both winml-ort and ort-genai reports use the same memory field names for shared concepts: RSS baseline, after-compile/load, after-inference, peak, model-load delta, inference/generation delta, and total delta; VRAM local/shared baseline, after-compile/load, after-inference, peak, model-load delta, inference/generation delta, and total delta.
With --runtime ort-genai, winml perf benchmarks the onnxruntime-genai decoder pipeline rather than a single session.run(). The JSON report uses a phase-based schema: load contains startup spans, requests contains one warmup or timed generation sample per request, aggregate summarizes timed requests only, memory contains optional RAM/VRAM deltas, and hw_monitor contains optional monitor output. The optional memory and hw_monitor top-level names match the classic winml-ort perf report; GenAI keeps load/requests/aggregate instead of classic latency_ms/throughput because generation has distinct prompt, first-token, and decode phases.
For model-ID auto-builds, the selected EP/device must be supported by the model's
bundle recipe; unsupported targets fail before export. --device auto retains
the concrete device selected by hardware detection. Prebuilt bundle directories
keep their existing runtime-override behavior and do not pass through this build
target validation.
GenAI metric definitions¶
| Core metric | JSON field(s) | Definition |
|---|---|---|
| Model Load Time | load.session_load_duration_ms, load.native_load_duration_ms |
session_load_duration_ms is the outer GenaiSession.load() wall-clock span. native_load_duration_ms is the onnxruntime-genai og.Config + og.Model + og.Tokenizer span. |
| Weight Upload Time | load.weight_upload_duration_ms, load.weight_upload_estimate_duration_ms |
Exact upload telemetry is null today because onnxruntime-genai does not expose it. The estimate is model_create_duration_ms and is labeled by weight_upload_estimate_source. |
| Cold Start Time | aggregate.cold_start_ttft_duration_ms, aggregate.cold_start_total_duration_ms |
Load plus the first request's TTFT, or load plus the first request's total request duration. |
| Warm Start Time / Latency | aggregate.request_duration_ms |
Timed-request full duration after warmups: template + tokenization + generator creation + model compute + sequence fetch + detokenization. |
| TTFT | requests[].model_ttft_duration_ms, requests[].request_ttft_duration_ms, aggregate.*ttft* |
Model TTFT is prefill + first-token compute. Request TTFT also includes template, tokenization, and generator creation. |
| Prefill TPS | requests[].prefill_tokens_per_second, aggregate.prefill_tokens_per_second |
Prompt tokens divided by prefill_duration_ms. |
| Decode TPS | requests[].steady_state_decode_tokens_per_second, aggregate.steady_state_decode_tokens_per_second |
Tokens after the first divided by the sum of per-token decode durations after the first. |
| RAM Usage | memory.rss_* |
Classic-compatible RSS fields such as rss_baseline_mb, rss_after_compile_mb, rss_after_inference_mb, rss_model_load_delta_mb, rss_inference_delta_mb, and rss_total_delta_mb. rss_checkpoint_peak_mb is the maximum of the sampled checkpoints, not a continuously sampled peak. Requires --memory. |
| VRAM Usage | memory.vram_* |
Adapter memory fields are emitted only when the effective GenAI route proves a specific accelerator adapter. Fields include baseline, after-compile, after-inference, load/inference/total deltas, and vram_*_checkpoint_peak_mb checkpoint maxima. Requires --memory. |
Examples¶
Basic benchmark on the best available device:
Device: npu
Precision: auto
Task: image-classification
Iterations: 100 (+ 10 warmup)
Batch Size: 1
Latency (ms)
Avg P50 P90 P95 P99 Min Max Std
2.14 2.11 2.38 2.51 2.79 1.97 3.04 0.12
Throughput: 467.29 samples/sec
Results saved to: ~/.cache/winml/perf/microsoft_resnet-50/2026-05-27T120000.json
Benchmark a pre-exported ONNX file on CPU with more iterations:
Benchmark a text model with an explicit task, targeting the NPU:
Benchmark with live hardware monitoring enabled:
Select one of multiple GPUs supported by the same EP:
$ winml sys
$ winml perf -m model.onnx --ep dml --device gpu --device-luid 0x00000000_0x00012C8B --monitor
Replace the example LUID with the adapter's value from winml sys. LUIDs are
local to the current Windows boot, not portable hardware IDs. The flag narrows
the resolved EP/device pair (and optional --ep name@source); it does not change
the EP/device auto-selection policy. It applies to ONNX, HuggingFace,
composite, and per-module runtime inference, including memory and hardware
monitoring. It does not pin the separate model-build compilation stage.
--device-luid cannot be combined with --ep-options device_id=..., even if
both identify the same adapter. This is rejected as a CLI usage error before
model or device resolution. Other provider options, such as performance tuning,
can still be used with --device-luid.
If multiple adapters match and --device-luid is omitted, perf warns and uses
the default ORT device unless --ep-options uniquely selects a concrete adapter.
This explicit selection suppresses the warning even when that adapter has no
LUID metadata. Non-selector options, or options whose adapter
binding cannot be established, do not suppress the warning.
When provider options resolve to another adapter of the same kind, that adapter
is used for the runtime binding, device identity, and monitoring. Options that
select a different device kind (for example, --device npu --ep qnn with
--ep-options backend_type=gpu) are rejected before the build; use the matching
--device instead so build and runtime agree.
If the bound adapter has no LUID, perf uses CPU/RAM monitoring mode rather than guessing another GPU/NPU. Adapter-specific utilization and VRAM sampling are disabled in that case. Separately labelled aggregate GPU telemetry may remain available; it is not attributed to the selected adapter.
The default GPU is selected by numeric DxgiHighPerformanceIndex from ORT
hardware metadata (0 first), then by LUID to break ties. Missing or invalid
indices sort after ranked GPUs, in LUID order. This is the same ordering used
by winml sys, restricted to the selected EP source's exposed devices rather
than all installed GPUs; it does not depend on ORT enumeration order.
An unavailable LUID fails instead of falling back to another adapter. Provider
options that redirect the binding away from the pinned adapter are rejected.
The requested pin is saved in
benchmark_info.device_luid in the single-model report.
Pass runtime EP provider options to tune the session (repeatable):
$ winml perf -m model.onnx --device npu \
--ep-options htp_performance_mode=burst \
--ep-options htp_graph_finalization_optimization_mode=3
Per-module benchmarking to find latency hot-spots across all attention blocks:
Benchmark with real inputs from a .npz file instead of random data:
Benchmark a Hugging Face model with dynamic export axes before measuring:
The archive must contain one array per model input, keyed by the input name — for example:
import numpy as np
np.savez("inputs.npz", pixel_values=np.zeros((4, 3, 224, 224), dtype=np.float32))
Array dtypes are cast to the model's expected dtype (with a warning) if they
differ, so .npz files saved with default integer/float widths still work.
When --op-tracing is combined with --input-data, the op trace runs on the
same real tensors as the latency benchmark (not on fresh random inputs). If the
traced graph's inputs don't match the provided data — for example a compiled
context model with different input names — the trace falls back to random inputs
and logs a warning.
Op-tracing results are included in the main benchmark JSON under
hw_monitor.ep_proof. The profiling CSV remains available as the raw trace
artifact; no separate _op_trace.json file is written.
Common pitfalls¶
- Warm-up too low on NPU. The first several inferences on an NPU EP can be significantly slower due to kernel compilation and caching. The default of 10 warm-up iterations is usually enough for vision models, but transformer models with many operators may need
--warmup 30or higher to reach steady-state latency. - Hidden third-party diagnostics. Normal
winml perfoutput suppresses noisy native warning-level diagnostics and Hugging Face download/progress chatter so benchmark results stay readable. Use-v/-vvor setWINMLCLI_SHOW_ALL_WARNINGS=1to show those warnings when debugging provider or Hub issues. --input-datakeys must match; dtypes are cast. The.npzkeys must equal the model's input names — a missing or unexpected key is a hard error (typo protection). Array dtypes are cast to the model's expected dtype with a warning (matching normal inference), so you don't have to hand-match widths..npyfiles are not supported — save named arrays as.npz. When--input-datais set,--batch-sizeand--shape-configare ignored (the tensors define their own shapes). It is also rejected for--modulemode,--runtime ort-genai, and composite (dual-encoder) models such as CLIP/SigLIP, where each sub-model has its own inputs that a single.npzcannot address.- Real data only binds if the export kept axes dynamic. When
-mis a HuggingFace model ID,perfexports it with default shapes (because--shape-config/--batch-sizeare ignored under--input-data). If that export baked in static shapes, ORT will reject differently-shaped--input-data. Use--dynamic-axes/symbolic--input-specsfor the Hugging Face build, or point-mat an ONNX file that already has dynamic axes. --shape-configis ignored when real or module inputs own the shape. It is ignored in--modulemode and when--input-datais set. The command prints a warning in both situations.- Random inputs do not represent real data distributions. Latency numbers are accurate, but memory access patterns may differ from production because the generated tensors are uniform random values. For memory-bandwidth-sensitive models this can understate real-world latency.
- Cross-device comparison. To compare performance across devices, run
winml perfseparately with different--devicevalues and compare the resulting JSON reports.
See also¶
- winml eval — measure accuracy after benchmarking
- winml build — build the quantized artifact that
perfbenchmarks - Load and export concept — how
--moduleper-instance benchmarking works - ONNX & Execution Providers — understand
--devicevs--ep