Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Combining & Stacking Scorers

Scorers are composable. Rather than building one complex scorer, combine small ones: aggregate several true/false scorers, invert a result, convert a float-scale score to a boolean with a threshold, or lift a message scorer to evaluate a whole conversation.

These wrappers are themselves scorers, so they plug into attacks and the batch scorer exactly like the leaf scorers on the True/False and Float-scale pages.

A wrapper must support every input condition through its children. Before scoring, it validates the whole tree and gives each child only its supported conditions, retaining objective context. For example, a Q&A/objective composite sends AnswerMatches to the Q&A judge and MatchesObjective to the objective judge. Direct leaves reject extra conditions. Missing criteria are errors, not skipped branches, even under OR.

The class hierarchy explains what each wrapper is. This diagram instead shows runtime composition: what each wrapper may contain. Solid arrows pass a scorer through scorer= or scorers=, while dashed arrows show the scorer base implemented by the resulting wrapper. An “any” input can therefore be a leaf scorer or an already composed wrapper with that base, which enables stacking.

TrueFalseCompositeScorer requires at least one TrueFalseScorer and combines their single results with AND, OR, or MAJORITY; TrueFalseInverterScorer accepts one TrueFalseScorer. FloatScaleThresholdScorer is the cross-kind adapter: it accepts one FloatScaleScorer and produces a TrueFalseScorer. FloatScaleFallbackScorer accepts two FloatScaleScorers and produces a FloatScaleScorer. TrueFalseFallbackScorer does the same for two TrueFalseScorers. Each fallback wrapper tries the primary first and calls the fallback only when the primary returns an undetermined score. These wrappers forward the same Scorable to their children, so each child must support that evidence kind.

create_conversation_scorer() accepts a true/false or float-scale scorer that supports text ContentScorable evidence. It returns a dynamic wrapper that remains the same scorer kind as its input.

An empty child result means that the scorer did not apply. A composite scorer ignores empty child results and aggregates the remaining results. It returns an empty list if every child result is empty. Inverter and threshold wrappers pass an empty result through unchanged. A conversation wrapper returns an empty result when it finds no applicable conversation evidence or its child returns no score. Any outer wrapper then applies the rules above.

Deprecated message-shaped calls remain on MessageScorer, but generic wrappers do not project those APIs from their children. Score wrappers through the canonical Scorable API.

For example, float-scale → conversation → threshold → inversion is supported; a generic Scorer outside those base types is not.

Found default environment files: ['./.pyrit/.env', './.pyrit/.env.local']
Loaded environment file: ./.pyrit/.env
Loaded environment file: ./.pyrit/.env.local
[pyrit:alembic] No new upgrade operations detected.

Composite true/false scorers

TrueFalseCompositeScorer aggregates several TrueFalseScorers into one result using an aggregator: AND, OR, or MAJORITY. Here two fast substring checks are combined.

[AND] both present -> True
[AND] one present  -> False

Inverting a true/false scorer

TrueFalseInverterScorer negates the wrapped scorer — handy when “no match” is the success condition (e.g. a refusal scorer where you want True when the model did not refuse).

[invert] 'bomb' absent -> True

Converting float-scale to true/false with a threshold

FloatScaleThresholdScorer wraps a FloatScaleScorer and returns True when the normalized score meets the threshold. This is the standard way to turn a severity score into a pass/fail success criterion. Below it wraps the local PlagiarismScorer.

[threshold] near-copy   -> True
[threshold] independent -> False

Routing abstentions to a fallback scorer

Both float-scale and true/false scorers can return ScoreStatus.UNDETERMINED. A fallback wrapper evaluates the same evidence with a second scorer only when the primary returns that status. A completed False or 0.0 is a valid judgment and does not trigger fallback. This lets a fast primary handle most inputs without calling an expensive second judge each time.

Both children must belong to the same result family, support the same condition types, and return exactly one score when applicable. The wrapper validates both children before scoring and passes each its supported expectation. If fallback runs, the results must refer to the same evidence and expectation; categories must match when both results supply them. An unreadable judgment can have no category labels. Configure equivalent criteria and, for float-scale scorers, comparable numeric meanings: matching types and categories does not prove that two rubrics measure the same thing.

With preconfigured scorers that meet these requirements:

from pyrit.score import FloatScaleFallbackScorer, TrueFalseFallbackScorer

harm_scorer = FloatScaleFallbackScorer(
    scorer=primary_harm_scorer,
    fallback_scorer=secondary_harm_scorer,
)
objective_scorer = TrueFalseFallbackScorer(
    scorer=primary_objective_scorer,
    fallback_scorer=secondary_objective_scorer,
)

A non-applicable primary ([]) returns [] without calling the fallback. A non-applicable fallback leaves the primary’s undetermined judgment in place. If both abstain, the result remains undetermined. Exceptions propagate; they are not treated as abstentions.

Multiple child results are rejected rather than paired by position or discarded. Aggregate them explicitly before using fallback if a single aggregate is meaningful.

Only the root score is persisted. The wrapper creates a new score without modifying either child result and retains their observation links. Metadata records resolved_by as "primary" or "fallback". Child metadata keys are prefixed with primary. and fallback., so duplicate keys and nested fallback details are not overwritten. When fallback is attempted, primary_rationale is retained and fallback_status records "complete", "undetermined", or "not_applicable". A fallback judgment also supplies fallback_rationale; the combined rationale explains both attempts.

Scoring a whole conversation

Some signals only emerge across turns — persuasion, gradual persona breaks, escalation. create_conversation_scorer() renders the conversation as text and passes that ContentScorable to a true/false or float-scale scorer. The returned scorer keeps the same result family as the scorer it wraps.

Pass it any one message from the conversation; its conversation_id is used to pull the full history from memory. Below we build a short conversation by hand and wrap a local SubStringScorer to flag a persona breach.

[conversation] persona breach across turns -> True

For a richer, real-world example, wrap a SelfAskLikertScorer with the BEHAVIOR_CHANGE_SCALE to measure how much a target’s behavior shifts over a multi-turn RedTeamingAttack — the wrapped float-scale scorer then rates the entire exchange.

Custom scorers

When the built-in templates don’t fit, the general self-ask scorers let you supply your own system prompt and JSON schema instead of writing a new class:

  • SelfAskGeneralTrueFalseScorer for boolean questions.

  • SelfAskGeneralFloatScaleScorer with a NumericRange for custom numeric ranges.

Both accept a system_prompt_format_string with {objective} placeholders and a rationale_output_key, so you can shape the scoring criteria without leaving Python.