Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Scorer Metrics

This demo will walk you through how to measure and gauge performance of PyRIT scoring configurations, both harm scorers and objective scorers.

Understanding Scorer Metrics

PyRIT has metrics for many scorers checked in. Before diving into how to retrieve or create metrics, it’s important to understand what they measure and how scorer identity determines which metrics apply.

Scorer Identity

Every scorer has a unique identity hash computed from its complete configuration:

  • Scorer type (e.g., SelfAskRefusalScorer)

  • System and user prompt templates

  • Target model information (model name)

  • Temperature and other generation parameters

  • Any scorer-specific configuration

This means changing any of these values creates a new scorer identity. The reason these are variables is because they might change performance—does changing the temperature increase or decrease accuracy? Metrics let you experiment and find out.

Metrics are stored and retrieved by this identity hash, so the same scorer configuration will always get the same cached metrics.

Found default environment files: ['./.pyrit/.env', './.pyrit/.env.local']
Loaded environment file: ./.pyrit/.env
Loaded environment file: ./.pyrit/.env.local
[pyrit:alembic] No new upgrade operations detected.
Scorer Identity:
  Type: SelfAskRefusalScorer
  System Prompt: **Important**: Your role is strictly that of an ev...
  Target Info: OpenAIChatTarget::99d992bd
  Identity Hash: 12cd53923d4290d9c97740cc75242610fbcfce508851d79116d09994b5ce9e6e

Objective Metrics

Objective scorers produce true/false outputs, usually whether an objective was met (e.g., did this response have instructions on “how to make a Molotov cocktail?”). We evaluate them using standard classification metrics by comparing model predictions against human-labeled ground truth.

  • Accuracy: Proportion of predictions matching human labels. Simple but can be misleading with imbalanced datasets.

  • Precision: Of all “true” predictions, how many were correct? High precision = few false positives.

  • Recall: Of all actual “true” cases, how many did we catch? High recall = few false negatives.

  • F1 Score: Harmonic mean of precision and recall. Balances both concerns.

  • Accuracy Standard Error: Statistical uncertainty in accuracy estimate, useful for confidence intervals.

Which metric matters most?

  • If false positives are costly (e.g., flagging safe content as harmful) → prioritize precision

  • If false negatives are costly (e.g., missing actual jailbreaks) → prioritize recall

  • For balanced scenarios → use F1 score

Harm Metrics

Harm scorers produce float scores (0.0-1.0) representing severity. Since these are continuous values, we use different metrics that capture how close the model’s scores are to human judgments.

Error Metrics:

  • Mean Absolute Error (MAE): Average absolute difference between model and human scores. An MAE of 0.15 means the model is off by 0.15 on average.

  • MAE Standard Error: Uncertainty in the MAE estimate.

  • Constant-Guess Baseline MAE (baseline_mean_absolute_error): The MAE of a scorer that ignores the response and always returns the dataset’s median human score, the best any constant can do on these labels. A scorer whose MAE is not below this has not beaten a constant guess on that dataset. Results recorded before this field existed report None.

Statistical Significance:

  • t-statistic: From a one-sample t-test. Positive = model scores higher than humans; negative = lower.

  • p-value: If small (e.g., < 0.05), the difference between model and human scores is statistically significant (not due to chance).

Inter-Rater Reliability (Krippendorff’s Alpha): Measures agreement between evaluators, ranging from -1.0 to 1.0:

  • 1.0: Perfect agreement

  • 0.8+: Strong agreement

  • 0.6-0.8: Moderate agreement

  • < 0.6: Weak agreement

Three alpha values are reported:

  • krippendorff_alpha_humans: Agreement among human evaluators (baseline quality of labels)

  • krippendorff_alpha_model: Agreement across multiple model scoring trials (model consistency)

  • krippendorff_alpha_combined: Overall agreement between humans and model

Error Split by Rater Agreement

When a gold set has more than one human rater, the harm metrics also report the mean absolute error separately for unanimous responses (every rater on the same side of contested_threshold, 0.5) and contested responses (the raters split across it, so the gold label rests on a 2-1 vote rather than a consensus). The aggregate MAE spends part of the scorer’s error budget on the contested rows, so a scorer can look strong overall while sitting near chance on exactly the responses humans found hard.

  • mean_absolute_error_unanimous and num_unanimous_responses

  • mean_absolute_error_contested and num_contested_responses

These are None for single-rater gold sets, where there is no disagreement to measure.

Retrieving Scorer Metrics

When scorer metrics are calculated with evaluate_async(), they can be saved to JSONL registry files and retrieved without re-running the evaluation. The PyRIT team has pre-computed metrics for common scorer configurations which you can access immediately.

Retrieving Cached Metrics for a Scorer

Use get_scorer_metrics() on any scorer instance to retrieve cached results matching its identity:

No cached metrics found for this scorer configuration.

Harm scorer metrics are retrieved similarly:

No cached metrics found for this scorer configuration.

Comparing All Scorer Configurations

The evaluation registry stores metrics for all tested scorer configurations. You can load all entries to compare which configurations perform best.

Use get_all_objective_metrics() or get_all_harm_metrics() to load evaluation results. These return ScorerMetricsWithIdentity objects with clean attribute access to both the scorer identity and its metrics.

Found 24 scorer configurations in the metrics file

Top 5 configurations by F1 Score:
--------------------------------------------------------------------------------

  📊 Scorer Information
    ▸ Scorer Identifier
      • Scorer Type: TrueFalseInverterScorer
      • scorer_type: true_false
      • score_aggregator: OR_
        └─ Composite of 1 scorer(s):
            • Scorer Type: SelfAskRefusalScorer
            • scorer_type: true_false
            • score_aggregator: OR_
            • model_name: gpt-4o-japan-nilfilter

    ▸ Performance Metrics
      • Accuracy: 89.37%
      • Accuracy Std Error: ±0.0155
      • F1 Score: 0.8918
      • Precision: 0.8782
      • Recall: 0.9058
      • Average Score Time: 0.59s

  📊 Scorer Information
    ▸ Scorer Identifier
      • Scorer Type: TrueFalseInverterScorer
      • scorer_type: true_false
      • score_aggregator: OR_
        └─ Composite of 1 scorer(s):
            • Scorer Type: SelfAskRefusalScorer
            • scorer_type: true_false
            • score_aggregator: OR_
            • model_name: gpt-5.4

    ▸ Performance Metrics
      Official evaluation has not been run yet for this specific configuration

  📊 Scorer Information
    ▸ Scorer Identifier
      • Scorer Type: TrueFalseInverterScorer
      • scorer_type: true_false
      • score_aggregator: OR_
        └─ Composite of 1 scorer(s):
            • Scorer Type: SelfAskRefusalScorer
            • scorer_type: true_false
            • score_aggregator: OR_
            • model_name: gpt-4o

    ▸ Performance Metrics
      Official evaluation has not been run yet for this specific configuration

  📊 Scorer Information
    ▸ Scorer Identifier
      • Scorer Type: TrueFalseInverterScorer
      • scorer_type: true_false
      • score_aggregator: OR_
        └─ Composite of 1 scorer(s):
            • Scorer Type: SelfAskRefusalScorer
            • scorer_type: true_false
            • score_aggregator: OR_
            • model_name: gpt-4o-japan-nilfilter

    ▸ Performance Metrics
      • Accuracy: 84.56%
      • Accuracy Std Error: ±0.0182
      • F1 Score: 0.8530
      • Precision: 0.7902
      • Recall: 0.9267
      • Average Score Time: 0.61s

  📊 Scorer Information
    ▸ Scorer Identifier
      • Scorer Type: TrueFalseInverterScorer
      • scorer_type: true_false
      • score_aggregator: OR_
        └─ Composite of 1 scorer(s):
            • Scorer Type: SelfAskRefusalScorer
            • scorer_type: true_false
            • score_aggregator: OR_
            • model_name: gpt-4o-unsafe

    ▸ Performance Metrics
      Official evaluation has not been run yet for this specific configuration

================================================================================
Best Accuracy:  89.37%
Best Precision: 0.990
Best Recall:    0.989
Fastest:        0.043 seconds
Slowest:        10.397 seconds

Similarly, you can look at the best harm scorers for a given category:

Found 4 harm scorer configurations for violence

Top configurations by Mean Absolute Error:
--------------------------------------------------------------------------------

  📊 Scorer Information
    ▸ Scorer Identifier
      • Scorer Type: SelfAskLikertScorer
      • scorer_type: float_scale
      • model_name: gpt-4o

    ▸ Performance Metrics
      • Mean Absolute Error: 0.1901
      • MAE Std Error: ±0.0229
      • Krippendorff Alpha (Combined): 0.8110
      • Krippendorff Alpha (Model): 0.9126
      • Average Score Time: 0.39s

  📊 Scorer Information
    ▸ Scorer Identifier
      • Scorer Type: SelfAskLikertScorer
      • scorer_type: float_scale
      • model_name: gpt-4o-japan-nilfilter

    ▸ Performance Metrics
      • Mean Absolute Error: 0.1901
      • MAE Std Error: ±0.0229
      • Krippendorff Alpha (Combined): 0.8110
      • Krippendorff Alpha (Model): 0.9126
      • Average Score Time: 0.39s

  📊 Scorer Information
    ▸ Scorer Identifier
      • Scorer Type: SelfAskLikertScorer
      • scorer_type: float_scale
      • model_name: gpt-4o-japan-nilfilter

    ▸ Performance Metrics
      • Mean Absolute Error: 0.1901
      • MAE Std Error: ±0.0229
      • Krippendorff Alpha (Combined): 0.8110
      • Krippendorff Alpha (Model): 0.9126
      • Average Score Time: 0.39s

  📊 Scorer Information
    ▸ Scorer Identifier
      • Scorer Type: AzureContentFilterScorer
      • scorer_type: float_scale

    ▸ Performance Metrics
      • Mean Absolute Error: 0.2437
      • MAE Std Error: ±0.0238
      • Krippendorff Alpha (Combined): 0.7754
      • Krippendorff Alpha (Model): 1.0000
      • Average Score Time: 0.72s

Creating Scorer Metrics

This section covers how to create new metrics by running evaluations against human-labeled datasets.

Caching and Skip Logic

When you call evaluate_async() on a scorer, the evaluation framework follows a smart caching strategy to avoid redundant work. It checks the metrics registry (a JSONL file) for an existing entry matching the scorer’s identity hash. The decision to skip or run evaluation depends on:

  1. No existing entry: Run the full evaluation

  2. Dataset version or harm definition version changed: Re-run and replace the old entry (assumes newer dataset/newer scoring criteria for harm is authoritative)

  3. Same version, sufficient trials: Skip if existing num_scorer_trials >= requested (existing metrics are good enough)

  4. Same version, fewer trials: Re-run with more trials and replace (higher fidelity needed)

During evaluation, the scorer processes each entry from human-labeled CSV dataset(s). For each assistant_response in the CSV, the scorer generates predictions which are compared against the human_score column(s). For objective scorers, this produces accuracy/precision/recall/F1 metrics. For harm scorers, it calculates MAE, t-statistics, and Krippendorff’s alpha.

Setting add_to_evaluation_results=False bypasses caching entirely—always running fresh evaluations without reading from or writing to the registry. This is useful for testing custom configurations without polluting the official metrics.

Running an Objective Evaluation

Call evaluate_async() on any scorer instance. The scorer’s identity (including system prompt, model, temperature) determines which cached results apply.

 Accuracy: 1.0

Running a Harm Evaluation

Metrics for harm category "exploits" created

Custom Evaluation Datasets

You can evaluate scorers against your own human-labeled datasets using ScorerEvalDatasetFiles:

# Configure custom datasets
scorer.evaluation_file_mapping = ScorerEvalDatasetFiles(
    # Glob patterns relative to SCORER_EVALS_PATH
    human_labeled_datasets_files=["my_datasets/*.csv"],
    # Where to store results
    result_file="my_datasets/evaluation_metrics.jsonl",
    # Required for harm scorers, ignored for objective scorers
    harm_category="violence",
)

CSV Human Evaluation Files

Many human scored dataset csv files are available in the pyrit/datasets/scorer_evals/ directory. These include datasets for refusal detection, hate speech, violence, and other harm categories. You can reference these as templates for creating your own evaluation datasets.

Required columns:

  • assistant_response: The text to score

  • objective: Required for objective scorers - the scoring question (e.g., “Did the response refuse?”)

  • harm_category: Required for harm scorers - the harm category (e.g., “hate_speech”, “violence”)

  • human_score or human_score_1, human_score_2, etc.: Ground truth labels from human raters

    • For objective scorers: 0 or 1 (converted to bool)

    • For harm scorers: 0.0-1.0 float values

  • data_type: Type of content (defaults to “text”)

Batch Evaluation with evaluate_scorers.py

While evaluate_async() runs evaluations for a single scorer, the evaluate_scorers.py script evaluates all registered scorers in bulk. This is useful for benchmarking after changing scorer prompts, adding new variants, or updating human-labeled datasets.

The script initializes PyRIT with ScorerInitializer (which registers all configured scorers), then runs evaluate_async() on each one. Results are saved to the JSONL registry files in pyrit/datasets/scorer_evals/.

Basic Usage

# Evaluate all registered scorers (long-running — can take hours)
python -m build_scripts.evaluate_scorers

# Evaluate only scorers with specific tags
python -m build_scripts.evaluate_scorers --tags refusal
python -m build_scripts.evaluate_scorers --tags refusal,default

# Control parallelism (default: 5, lower if hitting rate limits)
python -m build_scripts.evaluate_scorers --max-concurrency 3

Tags

ScorerInitializer applies tags to scorers during registration. These tags let you target specific subsets for evaluation:

  • refusal — The 4 standalone refusal scorer variants

  • default — All scorers registered by default

  • best_refusal_f1 — The refusal variant with the highest F1 (set dynamically from metrics)

  • best_objective_f1 — The objective scorer with the highest F1

When refusal scorer prompts or datasets change, the recommended workflow is:

Step 1: Evaluate refusal scorers first

python -m build_scripts.evaluate_scorers --tags refusal

This evaluates only the 4 refusal variants and writes results to refusal_scorer/refusal_metrics.jsonl. After this step, ScorerInitializer can determine which refusal variant has the best F1 and tag it as best_refusal_f1.

Step 2: Re-evaluate all scorers

python -m build_scripts.evaluate_scorers

On the next full run, ScorerInitializer reads the refusal metrics from Step 1, picks the best refusal variant, and uses it to build dependent scorers (e.g., TrueFalseInverterScorer wrapping the best refusal scorer). This ensures objective scorers that depend on refusal detection use the best-performing refusal prompt.

Scorers whose metrics are already up-to-date (same dataset version, sufficient trials) are automatically skipped, so re-running the full script is efficient.

Step 3: Commit updated metrics

git add pyrit/datasets/scorer_evals/
git commit -m "chore: update scorer metrics"

The updated JSONL files should be checked in so that ScorerInitializer can read them at runtime to select the best scorers.