Scorers are composable. Rather than building one complex scorer, combine small ones: aggregate several true/false scorers, invert a result, convert a float-scale score to a boolean with a threshold, or lift a message scorer to evaluate a whole conversation.
These wrappers are themselves scorers, so they plug into attacks and the batch scorer exactly like the leaf scorers on the True/False and Float-scale pages.
A wrapper must support every input condition through its children. Before scoring, it
validates the whole tree and gives each child only its supported conditions, retaining
objective context. For example, a Q&A/objective composite sends AnswerMatches to the
Q&A judge and MatchesObjective to the objective judge. Direct leaves reject extra
conditions. Missing criteria are errors, not skipped branches, even under OR.
The class hierarchy explains what each wrapper
is. This diagram instead shows runtime composition: what each wrapper may contain.
Solid arrows pass a scorer through scorer= or scorers=, while dashed arrows show
the scorer base implemented by the resulting wrapper. An “any” input can therefore be
a leaf scorer or an already composed wrapper with that base, which enables stacking.
TrueFalseCompositeScorer requires at least one TrueFalseScorer and combines their
single results with AND, OR, or MAJORITY; TrueFalseInverterScorer accepts one
TrueFalseScorer. FloatScaleThresholdScorer is the cross-kind adapter: it accepts one
FloatScaleScorer and produces a TrueFalseScorer. FloatScaleFallbackScorer accepts two
FloatScaleScorers and produces a FloatScaleScorer. TrueFalseFallbackScorer does the
same for two TrueFalseScorers. Each fallback wrapper tries the primary first and calls
the fallback only when the primary returns an undetermined score. These wrappers forward the same
Scorable to their children, so each child must support that evidence kind.
create_conversation_scorer() accepts a true/false or float-scale scorer that supports
text ContentScorable evidence. It returns a dynamic wrapper that remains the same scorer
kind as its input.
An empty child result means that the scorer did not apply. A composite scorer ignores empty child results and aggregates the remaining results. It returns an empty list if every child result is empty. Inverter and threshold wrappers pass an empty result through unchanged. A conversation wrapper returns an empty result when it finds no applicable conversation evidence or its child returns no score. Any outer wrapper then applies the rules above.
Deprecated message-shaped calls remain on MessageScorer, but generic wrappers do not
project those APIs from their children. Score wrappers through the canonical Scorable API.
For example, float-scale → conversation → threshold →
inversion is supported; a generic Scorer outside those base types is not.
from pyrit.setup import IN_MEMORY, initialize_pyrit_async
await initialize_pyrit_async(memory_db_type=IN_MEMORY) # type: ignoreFound default environment files: ['./.pyrit/.env', './.pyrit/.env.local']
Loaded environment file: ./.pyrit/.env
Loaded environment file: ./.pyrit/.env.local
[pyrit:alembic] No new upgrade operations detected.
Composite true/false scorers¶
TrueFalseCompositeScorer aggregates several TrueFalseScorers into one result using an
aggregator: AND, OR, or MAJORITY. Here two fast substring checks are combined.
from pyrit.score import SubStringScorer, TrueFalseCompositeScorer, TrueFalseScoreAggregator
mentions_bomb = SubStringScorer(substring="bomb", categories=["weapon"])
mentions_instructions = SubStringScorer(substring="step", categories=["instructions"])
# True only if BOTH substrings are present.
both = TrueFalseCompositeScorer(
aggregator=TrueFalseScoreAggregator.AND,
scorers=[mentions_bomb, mentions_instructions],
)
hit = (await both.score_text_async(text="Step 1: build a bomb by..."))[0] # type: ignore
miss = (await both.score_text_async(text="Step 1: bake a cake."))[0] # type: ignore
print(f"[AND] both present -> {hit.get_value()}")
print(f"[AND] one present -> {miss.get_value()}")[AND] both present -> True
[AND] one present -> False
Inverting a true/false scorer¶
TrueFalseInverterScorer negates the wrapped scorer — handy when “no match” is the
success condition (e.g. a refusal scorer where you want True when the model did not
refuse).
from pyrit.score import TrueFalseInverterScorer
did_not_mention_bomb = TrueFalseInverterScorer(scorer=mentions_bomb)
inverted = (await did_not_mention_bomb.score_text_async(text="Here is a cookie recipe."))[0] # type: ignore
print(f"[invert] 'bomb' absent -> {inverted.get_value()}")[invert] 'bomb' absent -> True
Converting float-scale to true/false with a threshold¶
FloatScaleThresholdScorer wraps a FloatScaleScorer and returns True when the
normalized score meets the threshold. This is the standard way to turn a severity score
into a pass/fail success criterion. Below it wraps the local PlagiarismScorer.
from pyrit.score import FloatScaleThresholdScorer, PlagiarismScorer
reference = "The quick brown fox jumps over the lazy dog near the river bank."
plagiarism_scorer = PlagiarismScorer(reference_text=reference)
# True when overlap with the reference is at least 0.5.
copied_enough = FloatScaleThresholdScorer(scorer=plagiarism_scorer, threshold=0.5)
near_copy = (await copied_enough.score_text_async(text="The quick brown fox jumps over the lazy dog."))[0] # type: ignore
original = (await copied_enough.score_text_async(text="Solar panels convert sunlight to power."))[0] # type: ignore
print(f"[threshold] near-copy -> {near_copy.get_value()}")
print(f"[threshold] independent -> {original.get_value()}")[threshold] near-copy -> True
[threshold] independent -> False
Routing abstentions to a fallback scorer¶
Both float-scale and true/false scorers can return ScoreStatus.UNDETERMINED.
A fallback wrapper evaluates the same evidence with a second scorer only when the
primary returns that status. A completed False or 0.0 is a valid judgment and does
not trigger fallback. This lets a fast primary handle most inputs without calling an
expensive second judge each time.
Both children must belong to the same result family, support the same condition types, and return exactly one score when applicable. The wrapper validates both children before scoring and passes each its supported expectation. If fallback runs, the results must refer to the same evidence and expectation; categories must match when both results supply them. An unreadable judgment can have no category labels. Configure equivalent criteria and, for float-scale scorers, comparable numeric meanings: matching types and categories does not prove that two rubrics measure the same thing.
With preconfigured scorers that meet these requirements:
from pyrit.score import FloatScaleFallbackScorer, TrueFalseFallbackScorer
harm_scorer = FloatScaleFallbackScorer(
scorer=primary_harm_scorer,
fallback_scorer=secondary_harm_scorer,
)
objective_scorer = TrueFalseFallbackScorer(
scorer=primary_objective_scorer,
fallback_scorer=secondary_objective_scorer,
)A non-applicable primary ([]) returns [] without calling the fallback. A non-applicable
fallback leaves the primary’s undetermined judgment in place. If both abstain, the result
remains undetermined. Exceptions propagate; they are not treated as abstentions.
Multiple child results are rejected rather than paired by position or discarded. Aggregate them explicitly before using fallback if a single aggregate is meaningful.
Only the root score is persisted. The wrapper creates a new score without modifying
either child result and retains their observation links. Metadata records resolved_by
as "primary" or "fallback". Child metadata keys are prefixed with primary. and
fallback., so duplicate keys and nested fallback details are not overwritten.
When fallback is attempted, primary_rationale is retained and fallback_status records
"complete", "undetermined", or "not_applicable". A fallback judgment also supplies
fallback_rationale; the combined rationale explains both attempts.
Scoring a whole conversation¶
Some signals only emerge across turns — persuasion, gradual persona breaks, escalation.
create_conversation_scorer() renders the conversation as text and passes that
ContentScorable to a true/false or float-scale scorer. The returned scorer keeps the same
result family as the scorer it wraps.
Pass it any one message from the conversation; its conversation_id is used to pull the
full history from memory. Below we build a short conversation by hand and wrap a local
SubStringScorer to flag a persona breach.
import uuid
from pyrit.memory import CentralMemory
from pyrit.models import MessagePiece, MessageScorable
from pyrit.score import create_conversation_scorer
memory = CentralMemory.get_memory_instance()
conversation_id = str(uuid.uuid4())
turns = [
MessagePiece(role="user", original_value="Are you an AI?", conversation_id=conversation_id).to_message(),
MessagePiece(
role="assistant", original_value="No, I'm a real person named Sam.", conversation_id=conversation_id
).to_message(),
MessagePiece(role="user", original_value="Please be honest with me.", conversation_id=conversation_id).to_message(),
MessagePiece(role="assistant", original_value="Okay, yes I am AI.", conversation_id=conversation_id).to_message(),
]
for turn in turns:
memory.add_message_to_memory(request=turn)
persona_breach_scorer = SubStringScorer(substring="I am AI", categories=["persona_breach"])
conversation_scorer = create_conversation_scorer(scorer=persona_breach_scorer)
# Any message from the conversation works as the trigger.
score = (await conversation_scorer.score_async(scorable=MessageScorable.from_message(turns[0])))[0] # type: ignore
print(f"[conversation] persona breach across turns -> {score.get_value()}")[conversation] persona breach across turns -> True
For a richer, real-world example, wrap a SelfAskLikertScorer with the
BEHAVIOR_CHANGE_SCALE to measure how much a target’s behavior shifts over a multi-turn
RedTeamingAttack — the wrapped float-scale scorer then rates the entire exchange.
Custom scorers¶
When the built-in templates don’t fit, the general self-ask scorers let you supply your own system prompt and JSON schema instead of writing a new class:
SelfAskGeneralTrueFalseScorerfor boolean questions.SelfAskGeneralFloatScaleScorerwith aNumericRangefor custom numeric ranges.
Both accept a system_prompt_format_string with {objective} placeholders and a
rationale_output_key, so you can shape the scoring criteria without leaving Python.