Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Trace-backed tool-call scoring

OtelToolCallScorer answers “Did these tools run?” It matches exact, case-sensitive names in execution spans, not claims in a response or model-proposed calls. A failed execution attempt counts; the scorer does not check tool success, arguments, order, or call count.

The example tool is named my_tool_call, implemented by the my_tool_call() function below. Its span records "gen_ai.tool.name": "my_tool_call". expects("my_tool_call") asks the scorer to find that exact tool name in the trace; "summarize" is a different tool name used to demonstrate a missing tool.

Supply trace IDs explicitly, or use stored message evidence from an attack. For messages, the scorer resolves request traces in the conversation through the scored response, without reading later turns.

The OpenTelemetry SDK is included with PyRIT. This walkthrough uses a real SDK provider, local capture, and PyRIT’s in-memory storage. It needs no model, service, credentials, or global provider changes.

Capture and score a real invocation

The caller gets the trace ID from its instrumented execution. The source selects only that trace and stores normalized evidence with the score.

Observed my_tool_call: True

Missing evidence is not a negative result

Capture starts incomplete. Observing a call proves invocation, but an absent tool is not false until coverage is explicitly complete.

This example disables span attribute string truncation, including limits set by environment variables. The exporter rejects finite or unknown length limits: shortened names cannot prove which tools ran, even with incomplete coverage.

This example has no background work and uses always-on sampling. Both spans have ended, so after checking export we can declare this controlled capture complete. A flush alone would not prove that an arbitrary remote trace is complete.

Missing tool before completion: undetermined
Missing tool after controlled completion: False

Re-match saved evidence with capture closed

Load the observation from PyRIT memory, stop capture, and change the expected name. Replay reads the same immutable snapshot; it does not call the trace client. Arguments and results are not retained. Scorer target response observations still have their stricter, original-expectation replay rules.

my_tool_call in saved evidence after capture is closed: True

Score a tool call through an attack

This local agent uses a real SDK provider and an in-process HTTP transport. No server, model, credentials, or global instrumentation is needed. HTTPTarget sends a fresh traceparent per request when you enable tracing, and PyRIT saves the same context on the request. The agent extracts that context before it runs a tool.

Here, the prompt "my_tool_call" tells the local agent to call the my_tool_call() tool. The prompt and tool name are the same only to keep the example simple. The scorer matches the tool name recorded in the span, not the prompt text.

The caller owns capture completeness. This controlled agent has no background work, sampling is disabled, and it checks export after its spans end. Only then does it mark a request trace complete. An HTTP response alone would not prove this for an arbitrary remote agent.

my_tool_call: success
no tool: failure
pending: undetermined

A composite receives the same message reference in each child. The message scorer reads the response, and the tool scorer reads linked execution evidence. This deterministic message scorer keeps the example offline; a normal LLM objective scorer can use the same composite path.

Message and tool evidence: success
Saved attack tool evidence: True

Configure tracing

HTTPTarget and HTTPXAPITarget disable tracing by default, because an arbitrary HTTP endpoint is not known to accept W3C trace context. Set trace_config=TargetTraceConfig(enabled=True) for an instrumented endpoint. Keep tracing disabled if you supply manual traceparent or tracestate headers.

A custom local target can pass TargetTraceConfig(tracer=provider.get_tracer(...)) to its PromptTarget constructor. This captures the target invocation with a caller-owned provider. PyRIT does not install a global provider or exporter.

For a remote agent, use the normal HTTP transport. The agent must accept W3C trace context and export its tool spans. Supply a TraceClient that can read those spans; sending a header does not create a trace-store connection. Provider-specific SDK targets are not changed by this example.

Use another trace source

A caller can provide a TraceClient that returns neutral spans for the requested trace IDs and reports coverage honestly. OtelTraceSource handles supported GenAI/OpenInference execution spans; matching remains backend-neutral. A client reports TraceQueryResult(available=False) when it knows no evidence can be retrieved, such as after retention expires. This produces an unavailable observation and an undetermined score. Pending capture stays available with incomplete coverage; an empty result alone does not mean unavailable. Optional empty call IDs are omitted; they do not prevent name-only scoring.

ObservationSource[TraceScorable] is the source contract for this scorer. Other sources can use the same protocol with their own scorable types. The caller owns instrumentation and trace retrieval. Automatic request correlation does not install a remote collector or a backend adapter.