Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

11. MessageNormalizer

MessageNormalizers convert PyRIT’s Message format into other formats that specific targets require. Different LLMs and APIs expect messages in different formats:

  • OpenAI-style APIs expect ChatMessage objects with role and content fields

  • HuggingFace models expect specific chat templates (ChatML, Llama, Mistral, etc.)

  • Some models don’t support system messages and need them merged into user messages

  • Attack components sometimes need conversation history as a formatted text string

The MessageNormalizer classes handle these conversions, making it easy to work with any target regardless of its expected input format.

Base Classes

There are two base normalizer types:

  • MessageListNormalizer[T]: Converts List[Message] → List[T] (e.g., to ChatMessage objects)

  • MessageStringNormalizer: Converts List[Message] → str (e.g., to ChatML format)

Some normalizers implement both interfaces.

Sample messages created:
  system: You are a helpful assistant....
  user: What is the capital of France?...
  assistant: The capital of France is Paris....
  user: What about Germany?...

ChatMessageNormalizer

The ChatMessageNormalizer converts Message objects to ChatMessage objects, which are the standard format for OpenAI chat-based API calls. It handles both single-part text messages and multipart messages (with images, audio, etc.).

Key features:

  • Single text pieces become simple string content

  • Multiple pieces become content arrays with type information

  • Supports use_developer_role=True for newer OpenAI models that use “developer” instead of “system”

ChatMessage output:
  Role: system, Content: You are a helpful assistant.
  Role: user, Content: What is the capital of France?
  Role: assistant, Content: The capital of France is Paris.
  Role: user, Content: What about Germany?
ChatMessage with developer role:
  Role: developer, Content: You are a helpful assistant.
  Role: user, Content: What is the capital of France?
  Role: assistant, Content: The capital of France is Paris.
  Role: user, Content: What about Germany?
JSON string output:
[
  {
    "role": "system",
    "content": "You are a helpful assistant."
  },
  {
    "role": "user",
    "content": "What is the capital of France?"
  },
  {
    "role": "assistant",
    "content": "The capital of France is Paris."
  },
  {
    "role": "user",
    "content": "What about Germany?"
  }
]

GenericSystemSquashNormalizer

Some models don’t support system messages. The GenericSystemSquashNormalizer merges the system message into the first user message using a standardized instruction format.

The format is:

### Instructions ###

{system_content}

######

{user_content}
Original message count: 4
Squashed message count: 3

First message after squashing:
### Instructions ###

You are a helpful assistant.

######

What is the capital of France?

ConversationContextNormalizer

The ConversationContextNormalizer formats conversation history as a turn-based text string. This is useful for:

  • Including conversation history in attack prompts

  • Logging and debugging conversations

  • Creating context strings for adversarial chat

The output format is:

Turn 1:
User: <content>
Assistant: <content>

Turn 2:
User: <content>
...
Conversation context format:
Turn 1:
user: What is the capital of France?
assistant: The capital of France is Paris.
Turn 2:
user: What about Germany?

TokenizerTemplateNormalizer

The TokenizerTemplateNormalizer uses HuggingFace tokenizer chat templates to format messages. This is essential for:

  • Local LLM inference with proper formatting

  • Matching the exact prompt format a model was trained with

  • Working with various open-source models

Using Model Aliases

For convenience, common models have aliases that automatically configure the normalizer:

AliasModelNotes
chatmlHuggingFaceH4/zephyr-7b-betaNo auth required
phi3microsoft/Phi-3-mini-4k-instructNo auth required
qwenQwen/Qwen2-7B-InstructNo auth required
llama3meta-llama/Meta-Llama-3-8B-InstructRequires HF token
gemmagoogle/gemma-7b-itRequires HF token, auto-squashes system
mistralmistralai/Mistral-7B-Instruct-v0.2Requires HF token
No HuggingFace token provided. Gated models may fail to load without authentication.
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
ChatML formatted output:
<|system|>
You are a helpful assistant.</s>
<|user|>
What is the capital of France?</s>
<|assistant|>
The capital of France is Paris.</s>
<|user|>
What about Germany?</s>
<|assistant|>

System Message Behavior

The TokenizerTemplateNormalizer supports different strategies for handling system messages:

  • keep: Pass system messages as-is (default)

  • squash: Merge system into first user message using GenericSystemSquashNormalizer

  • ignore: Drop system messages entirely

  • developer: Change system role to developer role (for newer OpenAI models)

No HuggingFace token provided. Gated models may fail to load without authentication.
ChatML with squashed system message:
<|user|>
### Instructions ###

You are a helpful assistant.

######

What is the capital of France?</s>
<|assistant|>
The capital of France is Paris.</s>
<|user|>
What about Germany?</s>
<|assistant|>

Using Custom Models

You can also use any HuggingFace model with a chat template by providing the full model name.

No HuggingFace token provided. Gated models may fail to load without authentication.
TinyLlama formatted output:
<|system|>
You are a helpful assistant.</s>
<|user|>
What is the capital of France?</s>
<|assistant|>
The capital of France is Paris.</s>
<|user|>
What about Germany?</s>
<|assistant|>

Creating Custom Normalizers

You can create custom normalizers by extending the base classes.

Markdown formatted output:
**System**: You are a helpful assistant.

**User**: What is the capital of France?

**Assistant**: The capital of France is Paris.

**User**: What about Germany?