Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Barge-In Attack (Streaming Audio)

BargeInAttack streams user audio to a RealtimeTarget and uses server-side voice-activity detection (VAD) to detect turn boundaries. When the user speaks while the assistant is still responding, server VAD cancels the in-flight response (barge-in). Interrupted turns are persisted with prompt_metadata["interrupted"] = True.

Audio converters are applied per turn after VAD commits. The raw audio drives interruption timing while the model responds to the converted version.

Note: Memory must be initialized via initialize_pyrit_async. See the Memory Configuration Guide.

Setup

BargeInAttack requires a RealtimeTarget with server_vad=True (or a ServerVadConfig for custom tuning).

Shared setup

Both sections use a pre-recorded 24 kHz mono PCM16 question about photosynthesis. The format matches what the OpenAI Realtime API expects. Any async generator yielding 24 kHz PCM16 bytes works as a chunk source (live mic, TTS, etc.).

Section 1: Single-turn streaming with a converter

Streams one user statement, applies a frequency-shift converter after VAD commits the turn, and gets the model’s response. Exercises the full pipeline (chunk push, convert-on-commit, item swap, response trigger, memory persistence) without barge-in.

Section 2: Barge-in (interrupting the assistant mid-response)

Plays the question twice with timing arranged so turn 2’s speech arrives during turn 1’s response. Server VAD detects the new speech, cancels turn 1’s response, and resolves it with interrupted=True.

Reading the barge-in output

After running the next cell, if barge-in fired successfully:

  • executed_turns: 2 (two VAD-detected user turns)

  • First assistant turn shows [INTERRUPTED] with a truncated transcript

  • Second assistant turn completes normally

If you don’t see [INTERRUPTED], decrease TURN1_RESPONSE_WAIT_S so turn 2’s audio arrives earlier in turn 1’s response window.

Alternate chunk sources

The chunk source is the main strategy hook:

  • Pre-recorded WAV (this notebook): most common starting point

  • TTS converter: generate audio from text prompts dynamically

  • Live microphone: use sounddevice or similar; yield what the mic produces

For feedback-driven attacks — for example, scoring each assistant turn and choosing to barge in with follow-up audio only when the response shows incomplete refusal — subclass BargeInAttack and override _perform_async to interleave turn observation with chunk generation.