Core concepts¶
This page explains the mental model behind Agora Workbench — the what and why. For what it is in one paragraph, see the home page; for step-by-step instructions, follow the links to the how-to guides below.
The core idea: code is the action surface¶
Most MCP integrations expose each capability as a separate MCP tool with a fixed JSON schema, and the agent invokes them one structured call at a time. Agora Workbench takes a different stance for the domain logic itself:
- A server exposes one primary MCP tool,
execute_{name}_code, that runs Python in a sandboxed, session-backed kernel. - Domain tools are ordinary Python functions injected into that kernel's namespace — they are not individual MCP tools. The agent calls them by writing code.
So the agent's loop is: discover what functions exist → write a snippet that calls and composes them → read the result. This "code as action" surface fits coding agents well and has concrete advantages:
- Composition is free. Loops, conditionals, intermediate variables, and chaining several tools happen in one round-trip instead of N tool calls.
- Context stays small. Only one tool schema sits in the agent's context, and intermediate data (DataFrames, arrays) stays server-side instead of being serialized back and forth through the conversation.
- Full expressiveness. The agent has all of Python plus your domain packages — not just the parameters a wrapper chose to expose.
The trade-off is that running code requires a sandbox — which is exactly what the per-server kernel environment provides.
→ See Tool pattern and Writing effective tools and skills.
Two kinds of server¶
| Server | Role |
|---|---|
CodeExecutionServer |
Runs a Python kernel, executes agent code, hosts domain tools and skills, manages sessions. The workhorse. |
ConnectorServer |
A stateless proxy (router / gateway / dispatcher) that aggregates and routes to several upstream servers — it has no kernel of its own. |
Both subclass BaseMCPServer and speak MCP over Streamable HTTP with
bearer-token auth.
→ See Server options and Server networks.
What a server actually exposes over MCP¶
A CodeExecutionServer registers a small, fixed set of MCP tools (names are
prefixed with the server name, e.g. chemistry):
| MCP tool | Purpose |
|---|---|
execute_{name}_code |
Run Python in the session kernel — the main action |
search_{name}_tools |
Discover domain tools and skills by natural-language query |
load_{name}_skill |
Load a skill's markdown instructions |
plan_{name}_workflow |
Navigate the state graph (only when tools declare state) |
{name}_list_sessions, {name}_inspect_session, {name}_close_session |
Inspect, list, or close sessions |
{name}_parallel_execute, {name}_check_batch |
Fan code out across many inputs |
{name}_send |
Send data to a destination — a peer server's kernel, blob storage, the user (browser download), or the local filesystem |
Sessions and state¶
The first execute_{name}_code call creates a session — a long-lived
kernel. Variables, imports, and loaded data persist across calls, so the
agent builds up state incrementally like a notebook rather than re-sending data
every turn. This persistence is the main thing a code-execution server gives you
over stateless tool calls.
Each execution response includes the execution session_id. To continue an
active session from another agent or a new MCP connection, pass that value as
execution_session_id to execute_{name}_code. This is distinct from the MCP
transport session ID; the server resumes only an existing session owned by the
caller, and rejects unknown or expired IDs rather than creating a new kernel.
→ See the Sessions API.
Three layers of guidance for the agent¶
- Tools — typed
ToolDefinitionfunctions, found viasearch_{name}_tools. - Skills — markdown playbooks for multi-step tasks, loaded via
load_{name}_skill. Preferred over manual planning when one fits the task. - State graph (optional) — tools can declare
requires/producesstates, andplan_{name}_workflowturns these into ordered paths. Registered only when tools carry state annotations. - Unified state graph (router only) — when a
RouterServeraggregates upstreams with state-annotated tools, it builds a composedStateGraphover all upstreams and exposesplan_{router_name}_workflow. Optionally,bridgesin the router config enable cross-server path queries (e.g., "graphormer output → battery simulation") without modifying individual upstream tool definitions.
→ See Skill pattern.
Working with data¶
- Asset provisioning — large, static files declared in
ServerConfig.assetsand fetched into the environment cache at startup. - Data catalog (optional) — a searchable index of datasets. When
configured, it adds the
search_data/get_artifact/list_domains/query_catalogtools (these names are not server-prefixed). - Asset tags — the agent embeds references like
<blob>id</blob>or<local>/path</local>as string literals; the server resolves them to localPathobjects before the code runs, handling download and auth for you.
→ See Working with data.
Sidecars: sharing a heavy resource¶
Because every session runs in its own kernel process, anything a tool loads is
loaded per session — fine for a small import, ruinous for a multi-gigabyte
model. A sidecar is a long-lived helper process the server launches at
startup so an expensive resource (typically a model) is loaded once per
container and shared by all sessions over loopback HTTP. Declare it in
ServerConfig.sidecars; the framework starts it, waits for its health check,
exports its URL to every kernel, and stops it on shutdown. A sidecar is a plain
internal service — not an MCP endpoint and invisible to the agent (contrast with
a ConnectorServer, which proxies other MCP servers).
→ See Sidecars.
Producing output: artifacts and publishing¶
Code writes user-facing files into the session output directory — addressed via
the injected agora_output("name") helper or the AGORA_OUTPUT_DIR variable.
To hand a file to the user, the agent calls {name}_publish_artifact with a
destination tag: <gui> (browser download, always available), <blob>, or
<local>.
Observability: the activity UI¶
Every server can stream what the agent is doing to the activity UI — a
standalone web dashboard (default http://127.0.0.1:8030) that shows code
executions, tool searches, skill loads, rendered plots, and published artifacts
in real time, grouped by session. Servers post events fire-and-forget; if the UI
isn't running, code execution is unaffected.
→ See Monitoring your servers.
Authentication¶
Auth is pluggable. Use no-op auth for local development (requests do not require a bearer token and any bearer token is accepted) and Entra ID for production. The server publishes RFC 9728 OAuth metadata so MCP clients can discover how to authenticate.
→ See Authentication options.
Where to go next¶
- Your First Server — build one end-to-end
- Options for making a CodeExecutionServer — every
ServerConfigknob - Attaching to your agent — connect a client and inject the workbench skill