Writing Eval Specs
A complete reference for writing eval.yaml specifications and task definitions.
Agent-ready workflow
Section titled “Agent-ready workflow”This guide is designed to be given directly to a coding agent. The agent should use the workflow below when creating or updating an eval suite, while developers can use the same checklist during review.
Required context
Section titled “Required context”Before editing, inspect:
- The target
SKILL.mdor.agent.md. - The existing
eval.yaml, task files, and fixtures, if present. schemas/eval.schema.jsonandschemas/task.schema.json.- The grader guide and existing nearby evals for repository conventions.
Treat the schemas as the source of truth for allowed fields. Do not invent YAML properties from examples or model memory.
Create an eval suite
Section titled “Create an eval suite”For a SKILL.md target:
# Scaffold from an existing skill.waza new eval <skill-name>
# Generate candidate coverage without writing files.waza suggest <path-to-SKILL.md> --dry-run
# Add generated tasks and fixtures to the scaffolded suite.waza suggest <path-to-SKILL.md> --apply --output-dir evals/<skill-name>waza new eval, waza suggest, and waza spec verify currently accept
SKILL.md targets only. For a .agent.md target, create eval.yaml, tasks,
and fixtures directly by following the custom agent guide
and the schemas linked above.
After scaffolding, the agent should:
- Map each
USE FOR,DO NOT USE FOR, parameter, and important behavioral requirement to at least one task. - Include positive, negative, boundary, and realistic integration cases.
- Use fixtures instead of relying on the developer’s working tree or live external state.
- Prefer deterministic graders (
text,file,behavior,action_sequence, orskill_invocation) when the expected behavior can be expressed mechanically. Use an LLM judge only for genuinely qualitative requirements. - Use
copilot-sdkfor real behavioral evaluation. Themockexecutor only verifies the harness and fixture plumbing; it does not evaluate agent quality.
Guide implementation with evals
Section titled “Guide implementation with evals”When an eval accompanies an implementation change, use a test-first loop:
- Add or update the task that expresses the required behavior.
- Run only that task and confirm that it fails for the expected reason.
- Change the skill, custom agent, or supporting implementation.
- Rerun the task until it passes without weakening the grader.
- Run related tasks, then the complete suite, to catch regressions.
Keep the task focused on observable behavior. Do not encode implementation details unless the requirement explicitly depends on a tool call, file change, or action sequence.
Update an existing eval suite
Section titled “Update an existing eval suite”When the target skill or agent changes:
- Read the diff and identify changed triggers, anti-triggers, parameters, tools, outputs, and failure behavior.
- For a
SKILL.mdtarget, runwaza spec verifyto find uncovered requirements. For a.agent.mdtarget, compare the tasks directly with the agent’s description, tools, handoffs, and behavioral instructions. - Add or update the smallest set of tasks needed to cover the changed behavior. Preserve stable task IDs unless the scenario itself is being replaced.
- Add a regression task for every bug fix or previously missed behavior.
- Keep unrelated tasks, fixtures, thresholds, and grader choices unchanged.
- If external tools are involved, use
mcp_mockswithschemaVersion: "1.1"or newer for deterministic CI coverage. Keep live-service testing in a separate integration suite.
File invariants
Section titled “File invariants”- Preserve the scaffolded or existing
schemaVersion. Do not copy a newer version unless the eval uses fields that require it.mcp_mocksrequires version1.1or newer;adversarialrequires version1.2or newer. - Every task requires non-empty
id,name, andinputs. inputsmust contain exactly one ofpromptorprompt_file.- Task IDs must be unique and stable; use lowercase kebab-case.
- Relative paths are resolved from the eval or task location described by the relevant schema field. Do not depend on the shell’s current directory.
- Fixtures must contain all state required by the scenario and must not contain secrets.
- Unknown MCP tools and unmatched mock calls should fail; do not add permissive fallback responses that hide missing fixtures.
- Before committing a behavioral suite, confirm
config.executoriscopilot-sdk. Never commitmockas a substitute for real evaluation. - Do not weaken thresholds merely to make a failing eval pass. Fix the skill, task, fixture, or grader that does not represent the intended behavior.
Validation ladder
Section titled “Validation ladder”Run the smallest relevant check first, then expand:
# SKILL.md only: inspect uncovered requirements without failing.waza spec verify <skill-path> <eval-path>
# SKILL.md only: enforce complete requirement coverage.waza spec verify <skill-path> <eval-path> --fail
# Exercise one changed task with the executor configured in eval.yaml.waza run <eval-path> --task "<task-id>" --no-cache -v
# Run the complete suite and save reviewable output.waza run <eval-path> --no-cache \ --output results.json \ --reporter junit:test-results.xmlSet config.executor: copilot-sdk before behavioral evaluation. Use
config.executor: mock only to check task loading, fixture copying, and grader
plumbing without a real agent session, then restore copilot-sdk before
committing.
For non-deterministic behavior, increase trials_per_task only after the
single-trial task is correct. Review failures individually; do not treat a
passing aggregate score as proof that every scenario is valid.
Delegation prompt
Section titled “Delegation prompt”Developers can give an agent this guide together with a focused request:
Create or update the Waza evals for <skill-or-agent-path>.
Follow the Agent-ready workflow in:https://microsoft.github.io/waza/guides/eval-yaml/
Requirements:- Treat schemas/eval.schema.json and schemas/task.schema.json as authoritative.- Preserve unrelated eval behavior and stable task IDs.- Cover positive triggers, negative triggers, parameters, edge cases, and the implementation behavior changed in <issue-or-diff>.- Use fixtures and mcp_mocks instead of live external state where possible.- For implementation work, add a focused failing task before changing behavior.- For SKILL.md targets, run waza spec verify.- Run a targeted task and the full relevant suite with config.executor set to copilot-sdk. Report the exact commands and results.eval.yaml Structure
Section titled “eval.yaml Structure”The evaluation spec defines the benchmark configuration, graders, and task files:
name: code-explainer-evaldescription: Evaluation suite for code-explainer skillskill: code-explainerschemaVersion: "1.0"version: "1.0"
config: trials_per_task: 1 timeout_seconds: 300 parallel: false executor: mock model: claude-sonnet-4.6
metrics: - name: accuracy weight: 1.0 threshold: 0.8
graders: - type: text name: explains_concepts config: regex_match: - "(?i)(function|logic|parameter)" - type: code name: has_output config: assertions: - "len(output) > 100"
tasks: - "tasks/*.yaml"Top-Level Fields
Section titled “Top-Level Fields”| Field | Type | Required | Description |
|---|---|---|---|
name |
string | ✓ | Eval suite name |
description |
string | ✗ | What the eval tests |
skill |
string | ✓ | Associated skill or custom agent name (SKILL.md or .agent.md) |
schemaVersion |
string | ✗ | Public schema version for this artifact; defaults to 1.0 |
version |
string | ✗ | Version number (e.g., “1.0”) |
inputs |
object | ✗ | Key-value map of global template variables (see Template Variables) |
tasks_from |
string | ✗ | Path to an external YAML file containing the task list |
hooks |
object | ✗ | Lifecycle hooks that run shell commands at specific points (see Hooks) |
mcp_mocks |
array | ✗ | Hermetic MCP server mocks for deterministic tool-call evals; requires schemaVersion: "1.1" |
command_mocks |
array | ✗ | Hermetic executable shims for CLI-dependent evals; requires schemaVersion: "1.3" or newer and executor: copilot-sdk |
adversarial |
object | ✗ | Built-in fault-injection packs consumed by waza adversarial --spec; requires schemaVersion: "1.2" |
baseline |
bool | ✗ | Mark this spec as a baseline for A/B comparison |
skill is required.
schemaVersion uses MAJOR.MINOR format. Omit it for legacy files that should default to 1.0; add it to new evals so future schema migrations are explicit. See Schema Changes for the compatibility policy.
Targeting Custom Agents
Section titled “Targeting Custom Agents”Waza supports evaluating VS Code custom agents (.agent.md files) alongside traditional SKILL.md-based skills. When you target an agent with a tools: field in its frontmatter, waza automatically injects a tool_constraint grader to validate that only the declared tools are called.
Specify an agent:
name: security-agent-evaldescription: Evaluate the security-reviewer custom agentskill: security-reviewer # Points to security-reviewer.agent.mdschemaVersion: "1.0"version: "1.0"
config: model: claude-sonnet-4.6Key differences:
- Use
skill: <name>to target either a skill or a custom agent - Waza discovers
.agent.mdfiles the same way asSKILL.md— in the current directory oragents/subdirectories - If both
SKILL.mdand.agent.mdexist in the same directory,SKILL.mdtakes priority - Custom agents can declare a
tools:field in frontmatter, which auto-injects atool_constraintgrader
Learn more: See the Evaluating Custom Agents guide for detailed examples and the auto-injected tool constraint behavior.
Config Section
Section titled “Config Section”The config block controls execution behavior:
config: trials_per_task: 1 # Run each task this many times timeout_seconds: 300 # Task timeout in seconds parallel: false # Run tasks sequentially (true = concurrent) workers: 0 # Auto-size parallel workers if parallel: true model: claude-sonnet-4.6 # Default model (override with --model) reasoning_effort: high # Optional: pin task and responder reasoning effort judge_model: gpt-4o # Model for LLM-as-judge graders (optional) judge_reasoning_effort: low # Optional: default effort for prompt graders executor: copilot-sdk # mock (local) or copilot-sdk (real API) inject_skill_body: true # Inject target SKILL.md/.agent.md body into the prompt instruction_files: - .github/instructions/project.instructions.md| Field | Type | Default | Description |
|---|---|---|---|
trials_per_task |
int | 1 | Number of times each task runs (for statistical analysis) |
timeout_seconds |
int | 300 | Task timeout in seconds |
first_event_timeout_seconds |
int | 0 (off) | Abort a run that produces no first event within N seconds (session-start hang); 0 disables |
parallel |
bool | false | Run tasks concurrently |
workers |
int | 0 | Number of parallel workers; 0 auto-sizes |
model |
string | required | Default model for tasks (override with --model flag) |
reasoning_effort |
string | model default | Copilot SDK task/responder effort: low, medium, high, xhigh, or max |
judge_model |
string | (same as model) |
Model for prompt-type graders (LLM-as-judge) |
judge_reasoning_effort |
string | judge default | Copilot SDK default effort for prompt graders, including continue_session graders; the grader can override it |
executor |
string | copilot-sdk |
Executor: mock (local, echoes task metadata and file content) or copilot-sdk (real API) |
max_attempts |
int | 0 | Maximum retry attempts per task on failure (0 = no retries) |
group_by |
string | — | Group results by a field (e.g., tags, task_id) |
fail_fast |
bool | false | Stop the entire run on first task failure |
skill_directories |
list[str] | [] |
Additional directories to search for skills |
instruction_files |
list[str] | [] |
Instruction files to apply to every task |
inject_skill_body |
bool | true | Inject the target SKILL.md or .agent.md body into the system prompt |
trigger_skill_routing |
bool | false | With inject_skill_body: false, add an eval-only routing instruction for trigger-precision runs |
disabled_skills |
list[str] | [] |
Skills to disable. Use ["*"] to disable all skills |
required_skills |
list[str] | [] |
Skills that must be available before running |
mcp_servers |
object | — | MCP server configurations for the evaluation. When tools is omitted or null for a server, Waza exposes all server tools by sending ["*"] to the Copilot SDK. Set tools to an explicit list, including [], to restrict or disable exposed tools. |
With an explicit reasoning effort, use a concrete model from waza models.
Waza checks the hosted runtime’s supported-effort metadata before session creation
and resumption, rejecting unsupported values rather than silently downgrading them.
Custom providers receive the effort directly because they are not covered by the
hosted model catalog. Prompt graders in eval, task, and checkpoint definitions may
override judge_reasoning_effort; agent effort stays eval-level.
Trigger-precision evals
Section titled “Trigger-precision evals”Set inject_skill_body: false when the eval is measuring whether the agent invokes a skill rather than whether it can complete the work after already seeing the skill body:
name: xyz-triggerdescription: Trigger-precision tasks for the xyz skillskill: xyzschemaVersion: "1.0"version: "1.0"
config: trials_per_task: 1 timeout_seconds: 300 parallel: false executor: copilot-sdk model: claude-sonnet-4.6 inject_skill_body: false trigger_skill_routing: true
metrics: - name: trigger_precision weight: 1.0 threshold: 0.8
tasks: - "tasks/trigger/*.yaml"trigger_skill_routing is optional and scoped to trigger-precision runs. It helps distinguish skill routing from general model fluency by telling the agent to invoke the target skill when the task is in scope, while still withholding the skill body. It has no effect unless inject_skill_body is false.
With this setting, Waza still discovers skills and passes their directories to the Copilot SDK, but it does not add the full target <skill_context> body or a synthetic <available_skills> summary. This lets behavior graders with required_tools or forbidden_tools and skill_invocation graders observe whether the skill tool was used. If you also set disabled_skills: ["*"], all skill loading is disabled and this setting has no effect.
Common Timeouts:
60— Quick tasks (single-file review, validation)300— Standard tasks (code explanation, analysis)600— Complex tasks (multi-file refactoring, design)
MCP Mock Servers
Section titled “MCP Mock Servers”Use top-level mcp_mocks to replace live MCP dependencies with deterministic local stdio servers. This keeps Copilot SDK evals hermetic in CI: no network listener, no port allocation, and no external service credentials. Waza automatically exposes every tool declared by a mock to the Copilot CLI, so no separate tools allowlist is needed. Because this is an additive eval schema field, set schemaVersion: "1.1" or newer.
name: issue-triage-evalskill: issue-triageschemaVersion: "1.1"version: "1.0"
config: executor: copilot-sdk model: claude-sonnet-4.6
mcp_mocks: - name: github tools: list_issues: description: Return matching issues for a repository input_schema: type: object properties: owner: { type: string } repo: { type: string } required: [owner, repo] responses: - match: owner: microsoft repo: waza return: issues: - number: 363 title: MCP server mocks for hermetic eval - match_regex: repo: "^waza-.*" return: issues: [] - match_schema: type: object required: [owner, repo] error: "No fixture for this repository"
tasks: - tasks/*.yamlResponse matching is evaluated in order. match requires exact full-argument equality, match_schema validates the call arguments against an inline JSON Schema, and match_regex applies regular expressions to individual argument fields. Unknown tools or calls that do not match any response fail loudly with an MCP tool error that points to the missing mock fixture.
You can keep larger fixtures in JSON files instead of inline YAML:
mcp_mocks: - name: github fixtures: fixtures/mcp/githubEach .json fixture can either contain a single tool definition (tool name from the filename) or a { "tools": { ... } } object with multiple tools.
Command Mocks
Section titled “Command Mocks”Use command_mocks when a skill shells out to a CLI such as Azure CLI (az), GitHub CLI (gh), or kubectl. Waza materializes temporary executable shims per task and prepends their directory to the Copilot runtime’s existing PATH, preserving the host path. Command mocks require executor: copilot-sdk and schemaVersion: "1.3" or newer.
schemaVersion: "1.3"config: executor: copilot-sdk model: claude-sonnet-4.6
command_mocks: - name: az expect_calls: 3 responses: - args: ["account", "show", "--output", "json"] stdout: subscriptionId: "00000000-0000-0000-0000-000000000000" name: Test Subscription - args_regex: ["group", "show", "--name", ".+"] fixture: fixtures/commands/az-group-show.json - args: ["deployment", "group", "create"] stderr: "deployment failed" exit_code: 1
tasks: - "tasks/*.yaml"args matches the entire argument vector, excluding the executable. args_regex applies a full-string regular expression to each argument at the same position. Responses are evaluated in order and the first matching response wins. stdout strings are emitted verbatim; structured values are JSON-encoded. fixture reads raw response bytes from a path relative to eval.yaml. Optional environment values must match exactly, and workdir matches a workspace-relative path. expect_calls requires an exact number of invocations for that executable; a mismatch is reported as a task error.
Task-level command_mocks replace the eval-level list. Set command_mocks: [] on a task to disable inherited mocks. A command that is configured but called with unmatched arguments fails closed with a diagnostic asking for a matching response fixture. Invocation records contain sanitized arguments, exit codes, and response indexes in runs[].command_invocations; graders can also read them from their context. Verbose runs print those sanitized details. Only declared executable names are intercepted, so command_mocks is not a command sandbox. Do not put real credentials in mock output.
Choosing a test boundary
Section titled “Choosing a test boundary”| Mechanism | Use it for |
|---|---|
command_mocks |
Commands the agent invokes through its shell, with deterministic process output and exit status. |
mcp_mocks |
Structured MCP tool calls and server responses. |
| Hooks | Setup, teardown, or validation around tasks/evals; hooks do not intercept agent commands. |
| Live integration tests | Verifying real CLI/service behavior, authentication, or infrastructure. Keep these separate from hermetic evals. |
Adversarial Packs
Section titled “Adversarial Packs”Use top-level adversarial to pin the built-in fault-injection packs and unsafe-outcome policy for waza adversarial --spec. This is additive in schemaVersion: "1.2" and is ignored by normal waza run executions.
schemaVersion: "1.2"adversarial: packs: - prompt-injection - scope-bypass on_unsafe_outcome: fail # or "warn"packs must name one or more built-in packs. In v0.38.0 those are prompt-injection and scope-bypass. on_unsafe_outcome: fail exits 2 when an unsafe outcome is observed; warn records the unsafe result but exits 0. See the Adversarial Harness guide for pack behavior and CI examples.
Graders Section
Section titled “Graders Section”Graders validate task outputs. Define once, reuse across tasks:
graders: - ref: github.com/waza-evals/fact#factuality@v1.0.0 name: factuality_strict weight: 2.0 config: threshold: 0.9
- type: text name: checks_logic weight: 2.0 config: regex_match: - "(?i)(function|variable|parameter)"
- type: code name: has_minimum_output config: assertions: - "len(output) > 100" - "'success' in output.lower()"
- type: text name: mentions_key_concepts config: contains: - "algorithm" - "optimization"Each grader accepts an optional weight (default 1.0) that controls its influence on the composite score. See Validators & Graders for details.
Remote grader presets can be referenced with ref instead of type. Refs use <host>/<owner>/<repo>[/path][#export]@<version> and resolve through waza get, which writes waza.lock with the pinned commit SHA and content digest. Local name, weight, top-level grader fields, and config values override the remote preset; nested config maps are deep-merged, while lists are replaced.
waza get eval.yamlwaza run eval.yamlwaza run requires the lockfile and cached module contents to be present for remote refs. It fails closed on missing locks, missing cache entries, or digest mismatches.
All graders return:
score: 0.0 to 1.0passed: booleanmessage: human-readable result
See the Validators & Graders guide for all 12 types and examples.
Mapping OpenAI Evals modelgraded YAML
Section titled “Mapping OpenAI Evals modelgraded YAML”OpenAI Evals modelgraded specs usually collapse into Waza’s prompt grader. The judge prompt carries the label semantics, while Waza handles execution and scoring.
| OpenAI Evals field | Waza equivalent | Notes |
|---|---|---|
prompt |
graders[].config.prompt |
Put the judging instructions directly in the prompt |
choice_strings |
prompt text | List the labels in the judge prompt; Waza’s prompt grader is binary, so the label choice becomes pass/fail guidance |
choice_scores |
prompt text | Encode the scoring rule in the judge prompt; use pairwise mode when the comparison is relative |
input_outputs |
tasks: entries |
Turn each example into one Waza task with its own inputs.prompt and expected checks |
eval_type: cot_classify |
type: prompt |
Use mode: independent for one-shot classification |
battle.yaml |
type: prompt + mode: pairwise |
Closest grader match for head-to-head comparison; waza compare is still the better run-level report |
Translation examples
Section titled “Translation examples”fact.yaml
Section titled “fact.yaml”OpenAI’s registry uses this pattern for fixed-choice factual classification. In Waza, keep the evaluation as a single prompt grader and turn each input/output row into a task:
graders: - type: prompt name: fact_check config: prompt: | You are checking a multiple-choice answer. Valid choices: A, B, C, D, E. Call set_waza_grade_pass only if the model's answer matches the correct choice. Otherwise call set_waza_grade_fail with a short reason. continue_session: false
tasks: - id: fact-001 name: fact-001 inputs: prompt: "Which answer is correct for the fact pattern?" expected: output_contains: - "B"closedqa.yaml
Section titled “closedqa.yaml”For closed-book QA, the judge prompt can encode the score mapping directly:
graders: - type: prompt name: closedqa_judge config: prompt: | Judge the answer against the reference. If the answer is fully correct, call set_waza_grade_pass. If it is partially correct or incorrect, call set_waza_grade_fail. Treat "Y" as 1.0 and "N" as 0.0 in your reasoning, but only emit pass/fail. model: claude-sonnet-4.5
tasks: - id: closedqa-001 name: closedqa-001 inputs: prompt: "Answer the question using the provided context." expected: output_contains: - "Y"battle.yaml
Section titled “battle.yaml”Battle-style comparisons are the one place where the mapping is not 1:1. The nearest Waza translation is a pairwise prompt grader, but the run-level comparison report is usually better expressed with waza compare:
config: baseline: true
graders: - type: prompt name: battle_judge config: mode: pairwise prompt: | Compare the two answers and decide which one is better. Call set_waza_grade_pass if the skill run wins. Call set_waza_grade_fail if the baseline run wins.
tasks: - id: battle-001 name: battle-001 inputs: prompt: "Compare these two solutions and pick the better one."Tasks Section
Section titled “Tasks Section”Tasks define individual test cases loaded from YAML files:
From Files
Section titled “From Files”Load tasks from YAML files in a directory:
tasks: - "tasks/*.yaml" # All YAML files in tasks/ - "tasks/basic/*.yaml" # Specific subdirectory - "tasks/advanced.yaml" # Single fileTask File Format
Section titled “Task File Format”Individual task files (e.g., tasks/basic-usage.yaml):
id: basic-usage-001name: Basic Usage - Python Functiondescription: Test that the skill explains a simple Python function correctly.
tags: - basic - happy-path
inputs: prompt: "Read sample.py and explain the function it contains." files: - path: sample.py
expected: output_contains: - "function" - "parameter" - "return" outcomes: - type: task_completed behavior: max_tool_calls: 5Task Fields
Section titled “Task Fields”| Field | Type | Description |
|---|---|---|
id |
string | Unique task identifier |
name |
string | Human-readable task name |
description |
string | What the task tests |
tags |
array | Tags for filtering (e.g., ["basic", "edge-case"]) |
inputs |
object | Test inputs (prompt, files) |
expected |
object | Validation rules and expected behavior |
skill_directories |
string[] | Skill directories for this task (overrides eval-level) |
instruction_files |
string[] | Instruction files for this task (adds to eval-level files) |
golden |
bool | Mark as a critical “golden” task — enforced by waza gate (must always pass) |
Inputs Section
Section titled “Inputs Section”inputs: prompt: "Your instruction to the agent" context: fixture: fixtures/demo # Optional: copy this fixture file/dir into the workspace files: - path: sample.py # Fixture file (relative to fixtures dir) content: | # Or inline content def hello(): print("Hello")inputs.files loads each path from the active fixtures directory (./fixtures
relative to the eval spec file by default) and copies it into the fresh task
workspace. It does not add file contents to prompt; tell the agent which
workspace file to read. Use content to define a file inline instead.
inputs.context.fixture is resolved relative to the eval spec directory. When it points to a directory, Waza copies that directory’s contents into the fresh task workspace before the agent runs.
Loading prompts from a file
Section titled “Loading prompts from a file”Use prompt_file instead of prompt to load the prompt text from an external file.
The path is resolved relative to the task YAML file’s directory.
inputs: prompt_file: prompts/review-instructions.md files: - path: sample.pyThis is useful when prompts are long, shared across tasks, or maintained separately.
You must specify either prompt or prompt_file, but not both.
Follow-up Prompts
Section titled “Follow-up Prompts”Use follow_up_prompts to send additional messages after the initial prompt. Each follow-up reuses the same session and workspace, so file changes and conversation history persist across turns.
inputs: prompt: "Create a Python function that reads a CSV file" follow_up_prompts: - "Add error handling for missing files" - "Write unit tests for the function"This is useful for evaluating multi-turn conversations where each step builds on the previous one. Graders run only after all prompts (initial + follow-ups) have completed, so the final output reflects the full conversation.
Responder (Interactive Skills)
Section titled “Responder (Interactive Skills)”For skills that ask follow-up questions, configure a responder — an LLM that plays the user and answers the skill’s questions. It is mutually exclusive with follow_up_prompts.
inputs: prompt: "Add a new agent to my application" responder: model: gpt-4o # optional; defaults to config.model instructions: | The agent you want is "research-agent" with system instructions "Search the web and summarise findings", tools web_search + url_fetch, and no handoffs. Answer the skill's questions consistently with this. If you genuinely can't infer an answer, abstain. max_followups: 8After each agent turn the responder either replies (the answer is sent back, continuing the conversation), stops (the agent is done), or abstains — which fails the run with a distinct abstained outcome, signalling the brief is too vague. If max_followups is reached while the agent is still asking questions, the loop stops with outcome cap_exhausted and graders evaluate the final state. Each task carries its own responder, so the same skill can be tested against several target configurations.
Per-Turn Checkpoints
Section titled “Per-Turn Checkpoints”By default graders only run once, against the final state of a multi-turn conversation. For long conversations you can run graders at specific turn boundaries using a top-level checkpoints: list:
checkpoints: - after_turn: 1 graders: - type: text contains: ["analyzing", "files"] - after_turn: 2 on_failure: stop # abort the run if this checkpoint fails graders: - type: tool_calls required: ["read_file"]Each checkpoint accepts:
| Field | Type | Description |
|---|---|---|
after_turn |
int | 1-based turn number this checkpoint runs after (initial prompt is turn 1). |
graders |
array | Inline graders, same schema as the task-level graders: / eval-level graders: field. |
on_failure |
string | continue (default) or stop — abort remaining turns when this checkpoint fails. |
Outcomes are recorded per-checkpoint on results.json under checkpoints[], alongside the final validations. waza gate still uses final-pass status. Available with schemaVersion: "1.1" and above (additive — older 1.0 files load unchanged).
When a task includes workspace files, name the file in the prompt so the agent knows to read it:
inputs: prompt: "Read sample.py and explain this code." files: - path: sample.pyExpected Section
Section titled “Expected Section”expected: # Strings that must appear in output output_contains: - "function" - "parameter"
# Output must NOT contain these output_not_contains: - "error" - "failed"
# At least one of these must appear (flexible matching) output_contains_any: - "recursion" - "iteration" - "loop"
# Task outcomes outcomes: - type: task_completed - type: tool_called tool_name: code_analyzer
# Behavioral constraints behavior: max_tool_calls: 5 max_tokens: 4096output_contains vs output_contains_any
Section titled “output_contains vs output_contains_any”output_contains— ALL listed strings must appear (AND logic). Use for required content.output_contains_any— At least ONE listed string must appear (OR logic). Use when the agent may express concepts in different ways.
All checks are case-insensitive.
Fixture Isolation
Section titled “Fixture Isolation”Fixtures are test files (code, documents, data) that tasks reference.
Important: Each task gets a fresh temp workspace with fixtures copied in. Original fixtures are never modified.
Using Fixtures
Section titled “Using Fixtures”Create a fixtures/ directory:
evals/code-explainer/├── eval.yaml├── tasks/│ └── basic-usage.yaml└── fixtures/ ├── sample.py ├── complex.py └── README.mdReference in tasks:
inputs: prompt: "Read sample.py and analyze it." files: - path: sample.pyInstruction Files
Section titled “Instruction Files”Use instruction_files for repository or task-specific *.instructions.md guidance:
config: instruction_files: - .github/instructions/project.instructions.mdinstruction_files: - .github/instructions/review.instructions.mdinputs: prompt: "Review this change" files: - path: sample.pyInstruction files are resolved from the active fixtures/context directory, copied into each fresh temp workspace, and appended to the agent system message with path labels. Eval-level files apply to every task; task-level files are added for that task. Paths must be relative and cannot use directory traversal.
Directory Structure
Section titled “Directory Structure”# Project modeevals/└── code-explainer/ ├── eval.yaml ├── tasks/ │ ├── basic-usage.yaml │ ├── edge-case.yaml │ └── should-not-trigger.yaml └── fixtures/ ├── sample.py ├── complex.py └── nested/ └── module.pySpecify context directory when running:
waza run eval.yaml --context-dir evals/code-explainer/fixturesOr use relative paths in eval.yaml if fixtures are adjacent.
Multi-Model Comparison
Section titled “Multi-Model Comparison”Run the same eval against multiple models:
# Run with gpt-4owaza run eval.yaml --model gpt-4o -o gpt4.json
# Run with Claudewaza run eval.yaml --model claude-sonnet-4.6 -o sonnet.json
# Compare resultswaza compare gpt4.json sonnet.jsonOverride the default model in eval.yaml:
waza run eval.yaml --model gpt-4o # Overrides config.modelFiltering and Parallel Execution
Section titled “Filtering and Parallel Execution”Filter by Task Name
Section titled “Filter by Task Name”waza run eval.yaml --task "basic*" --task "edge*"Filter by Tags
Section titled “Filter by Tags”waza run eval.yaml --tags "happy-path"Parallel Execution
Section titled “Parallel Execution”# Run tasks concurrently with auto-sized workerswaza run eval.yaml --parallelSaving Results
Section titled “Saving Results”Save eval results for later analysis or comparison:
waza run eval.yaml -o results.jsonOutput format:
{ "name": "code-explainer-eval", "model": "claude-sonnet-4.6", "pass_rate": 0.8, "tasks": [ { "id": "basic-001", "name": "Basic Usage", "passed": true, "graders": [ { "name": "checks_logic", "passed": true, "score": 1.0 } ] } ]}Caching
Section titled “Caching”For iterative testing, cache results:
waza run eval.yaml --cache --cache-dir .waza-cacheOnly tasks with changed inputs/config re-run.
Common Patterns
Section titled “Common Patterns”Simple Validation
Section titled “Simple Validation”graders: - type: text name: format_check config: regex_match: - "^[A-Z].*\\.$" # Sentence starting with capital, ending with period
tasks: - "tasks/format/*.yaml"Multi-Criteria Scoring
Section titled “Multi-Criteria Scoring”graders: - type: code name: completeness config: assertions: - "len(output) > 500" - "'function' in output" - "'parameter' in output"
tasks: - "tasks/completeness/*.yaml"Behavioral Constraints
Section titled “Behavioral Constraints”Behavioral constraints are defined in individual task YAML files:
id: efficient-001name: Efficiency testinputs: prompt: "Refactor this code"expected: behavior: max_tool_calls: 3 # Efficient max_tokens: 1000 # Concise max_response_time_ms: 30000 # Must complete within 30 seconds required_tools: # Must use these tools - grep - edit forbidden_tools: # Must NOT use these tools - rm| Field | Type | Description |
|---|---|---|
max_tool_calls |
int | Maximum number of tool invocations allowed |
max_iterations |
int | Maximum number of conversation rounds (turns) |
max_tokens |
int | Maximum tokens in the response |
max_response_time_ms |
int | Maximum wall-clock execution time in milliseconds |
required_tools |
string[] | Tools the agent must use during the task |
forbidden_tools |
string[] | Tools the agent must NOT use during the task |
Each constraint that is set (non-zero / non-empty) contributes equally to the behavior efficiency score. If all constraints pass, the score is 1.0; each failure reduces it proportionally.
Lifecycle hooks run shell commands at specific points during an evaluation. Use them for setup, teardown, or validation.
hooks: before_run: - command: "npm install" working_directory: "./fixtures" error_on_fail: true after_run: - command: "bash cleanup.sh" before_task: - command: "echo Starting task" after_task: - command: "bash collect-metrics.sh"| Hook | When it runs |
|---|---|
before_run |
Once, before the entire evaluation starts |
after_run |
Once, after all tasks complete |
before_task |
Before each individual task |
after_task |
After each individual task |
Each hook entry:
| Field | Type | Default | Description |
|---|---|---|---|
command |
string | (required) | Shell command to execute |
working_directory |
string | . |
Working directory for the command |
exit_codes |
list[int] | [0] |
Acceptable exit codes |
error_on_fail |
bool | false | Abort the run if this hook fails |
Template Variables
Section titled “Template Variables”Use the inputs field to define global template variables that are substituted into task prompts:
inputs: language: python framework: fastapi
tasks: - "tasks/scaffold/*.yaml"Fixture files are copied into the task workspace, not inlined in the prompt:
inputs: prompt: "Read sample.py and explain this code." files: - path: sample.pyExternal Task Lists
Section titled “External Task Lists”Use tasks_from to load task definitions from a separate YAML file:
name: shared-evaltasks_from: shared-tasks.yaml
config: trials_per_task: 3 model: claude-sonnet-4.6This is useful when multiple eval specs share the same task set but differ in config or graders.
Best Practices
Section titled “Best Practices”- Clear task descriptions — Future reviewers should understand what’s being tested
- Realistic validators — Don’t over-specify. A few key checks beat 20 strict rules
- Fixture diversity — Include basic, edge case, and negative test fixtures
- Tag your tasks — Makes filtering and analysis easier
- Use timeout appropriately — Too short = false failures, too long = slow tests
- Reuse graders — Define once, apply across multiple tasks
- Version your evals — Track improvements with version numbers
Next Steps
Section titled “Next Steps”- Validators & Graders — Reference for all grader types
- Web Dashboard — Explore results interactively
- CLI Reference — All commands and flags