Skip to content

Scope

Evaluate the agentic experience of your products

Run one scenario across GitHub Copilot, Claude Code, and VS Code. Score every run against criteria you define, and see how your agentic experience holds up across all of them — at scale.

How it works

One task, every agent, measured the same way

Scope submits your task to each agent, scores every run against the criteria you define, and lines the results up side by side.

01

Submit a task

Hand one coding task to GitHub Copilot, Claude Code, or VS Code with the Copilot driver extension — with the profile you want.

02

Score against criteria

Each run is evaluated against your criteria DAG. Gate child criteria on their parents, or leave them as independent roots.

03

Compare trajectories

See not just whether it worked, but how each agent got there — the tools it reached for and where it got stuck.

What you can do

Not “did it work?” — how did it get there?

Scope drives existing agents against tasks you control, then captures what they asked for, what tools they reached for, and where they got stuck.

01

Submit a coding task

Hand a task to GitHub Copilot, Claude Code, or VS Code with the Copilot driver extension.

02

Save reusable profiles

Bundle worker, model, version, MCP servers, skills, and extensions to re-run a setup consistently.

03

Define criteria as a DAG

Describe what “good” means. Gate child criteria on their parents, or leave them as independent roots.

04

Watch runs live

Stream logs straight from the worker as each run executes, attempt by attempt.

05

Slice with prompt features

Scope detects characteristics on your prompts so you can compare across heterogeneous tasks.

06

Automate everything

Submit runs, manage profiles and criteria, and fetch results through the REST API or the CLI.

Three ways to drive it

Portal, REST API, or CLI

Point & click

Portal

Submit requests, configure profiles and criteria, and watch runs stream — all from the browser.

Submit from the Portal →
Integrate

REST API

Wire Scope into your pipelines and tooling with a single authenticated request.

Read the API guide →
Terminal

CLI

Run benchmarks from the terminal with real-time log streaming.

Install the CLI →
Documentation

Explore the docs

Your first run

Run a task end to end from the Portal in a few minutes. Get started.

Submit via REST API

Automate request submission and integrate Scope into your tools. API guide.

Reference

Endpoint reference, configuration schemas, and worker capabilities. Browse the reference.

Stop guessing which agent is better. Measure it.

Spin up your first benchmark in minutes — from the Portal, the API, or the CLI.