Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

HackAPrompt Target

HackAPrompt is a prompt injection competition. Each challenge gives you a model to break and an intent to reach, and a panel of judges run by the platform decides whether a session reached it. HackAPromptTarget sends prompts to one challenge and can ask that judge panel for its verdict, so the challenges can be attacked with PyRIT attacks and scorers instead of the website’s chat box.

Only the first cell below runs without credentials. The two that talk to the platform need a signed-in HackAPrompt session, so their output is not checked in and this notebook is skipped when the documentation is regenerated.

Authentication

HackAPrompt has no API keys. Requests are authenticated with the cookies of a signed-in browser session, and a challenge session is identified by a session id that the website creates when you open a challenge. Both values are copied out of the browser once and expire when the browser session does.

  1. Sign in at hackaprompt.com and open the challenge you want to attack.

  2. Open the browser’s developer tools (F12) and switch to the Network tab.

  3. Send any prompt in the challenge’s chat box and select the resulting chat request.

  4. The request payload holds session_id, challenge_slug and competition_slug.

  5. The request headers hold the Cookie header. Copy the whole header value; it holds the sb-<project>-auth-token.* cookies that identify you.

Put the two secrets in your .env file, where the target picks them up:

HACK_A_PROMPT_SESSION_ID="a4d3e38c-..."
HACK_A_PROMPT_COOKIE="sb-<project>-auth-token.0=...; sb-<project>-auth-token.1=..."

They can also be passed to the constructor as session_id and cookie — useful when you refresh them from a browser in the same script. Neither ends up in the target identifier that PyRIT writes to memory.

Picking a challenge

HackAPromptChallenge lists the text challenges of the CBRNE practice track, the live twin of the CBRNE competition track that closed on June 19th. Picking a member fills in both slugs the API needs, and carries the rest of what the platform publishes about the challenge: the intents the judges grade against, the models the prompt is run against, and whether the challenge takes a single prompt only.

BACTERIAL_BASICS                  basic_challenge_cbrne_practice      conversation not published
AUTONOMOUS_ATOMICS                basic_uranium_munitions_practice    conversation not published
KOLOKOL_KALINKA                   basic_kolokol_agent_practice        conversation not published
PATHOGEN_PILGRIMAGE               multi_pathogen_pilgrimage_practice  conversation crisp_meadow, tranquil_lake, optimistic_peak, urban_lantern
FEATHERED_FALLOUT                 multi_feathered_fallout_practice    conversation crisp_meadow, tranquil_lake, optimistic_peak, urban_lantern
CRISPR_COLLAPSE                   multi_crispr_collapse_practice      conversation crisp_meadow, tranquil_lake, optimistic_peak, urban_lantern
WEST_NILE_WOES                    opam_nile_practice                  one prompt   quick_tree, brave_journey, optimistic_peak, urban_lantern
RICIN_RAMPAGE                     opam_ricin_practice                 one prompt   quick_tree, brave_journey, optimistic_peak, urban_lantern
PUFFERFISH_PARTYFOUL              opam_puff_practice                  one prompt   quick_tree, brave_journey, optimistic_peak, urban_lantern
TRIPLE_TOXIN_THREAT               triple_toxin_threat_practice        one prompt   brave_journey, urban_lantern
PATHOGEN_POLY_PROBLEM             pathogen_poly_problem_practice      one prompt   brave_journey, urban_lantern
MISCELLANEOUS_MALICIOUS_MENAGERIE misc_malicious_menagerie_practice   one prompt   brave_journey, urban_lantern

Sending a prompt

One target drives one HackAPrompt challenge session: the platform keeps the transcript on its side, keyed by the session id, and the judges grade all of it together. Build a new target — or assign a fresh session_id copied from a freshly opened challenge page — for every attack run that should be graded on its own.

A challenge the platform does not flag as one-shot is a conversation, so multi-turn attacks such as MultiPromptSendingAttack and ChunkedRequestAttack work against it; a one-shot challenge rejects a second turn.

Attacks that rewrite the history as they go — CrescendoAttack among them — do not work here, and the target does not pretend otherwise. Their TARGET_REQUIREMENTS ask for native EDITABLE_HISTORY, which this target does not declare, so AttackStrategy.__init__() raises ValueError before anything is sent. The transcript belongs to the platform: it is keyed by the session id, it is what the judges grade, and nothing sent from here can edit or truncate it. Declaring the capability to get past the check would only move the failure somewhere harder to read.

Asking the judges

The attack above runs without a scorer, so PyRIT reports its outcome as undetermined. On HackAPrompt it is the platform’s own judge panel that decides: check_challenge_async submits the session to it and returns its verdict unchanged, one entry per judge plus the points the session earned. The judges grade the whole session, so call it after the attack has run.

Other challenges

HackAPrompt has far more challenges than the CBRNE practice track, and new ones appear with every competition. Any of them can be attacked by passing its slug together with the slug of the competition it belongs to — both are in the payload of the chat request, and in the address of the challenge page. This is also how the challenges of the closed CBRNE competition track are reached.

target = HackAPromptTarget(challenge_slug="jokebot_goes_to_therapy", competition_slug="dougdoug")