Add a data catalog to your server¶
Use the data-lake API when an agent needs to discover data by metadata and then load the selected artifact into a code-execution session.
If you only need to copy fixed files into the server at startup, use asset provisioning. If you only need to return generated files to a user or storage location, use publishers. Neither requires a catalog.
Choose the workflow¶
| Goal | Start with |
|---|---|
| Let agents search existing local files | A read-only catalog with discovery: scan |
| Expose only an approved list of files | A read-only catalog with discovery: manifest |
| Register, upload, remove, or promote durable artifacts | An authoritative manifest plus an authorized managed writer |
| Use Azure Blob Storage | The same catalog model with the azure extra |
| Upgrade an existing 0.2.x deployment | 0.3 migration guide |
For a first deployment, start with keyword-only local search. It requires no cloud account, embedding service, vector extension, or credentials.
The five parts¶
The components have separate jobs:
- A source is the local directory or Blob prefix containing data.
- A catalog configuration says how files are discovered and describes their default metadata.
- The SQLite index makes catalog metadata searchable. It is rebuildable; it is not the source of truth for managed artifacts.
CatalogIntegrationadds policy-aware discovery and opaque load references to aCodeExecutionServer.- A session's data manager resolves a selected reference and caches its bytes locally for tool or code execution.
Publishing an output and registering it in a catalog are different operations. A publisher moves bytes. A managed writer performs an authorized, versioned catalog mutation.
Try a local catalog¶
Install the base package:
Create a directory with synthetic data:
Create, validate, and index a scan catalog:
uv run agora-workbench-data-lake init \
--config catalog.yaml \
--source ./data \
--source-id local-data \
--discovery scan
uv run agora-workbench-data-lake validate --config catalog.yaml
uv run agora-workbench-data-lake refresh \
--config catalog.yaml \
--database catalog.db
uv run agora-workbench-data-lake search weather \
--config catalog.yaml \
--database catalog.db
The search result should include weather.csv. At this point you have tested
the source policy, configuration, index, and search behavior without starting
an MCP server.
Scan mode exposes the directory
Scan discovery catalogs every eligible non-hidden file below the source root. Use it only when the whole directory is approved for discovery. For selective disclosure, use an authoritative manifest.
Add the catalog to a server¶
Add CatalogIntegration to an existing CodeExecutionServer:
from agora_workbench.code_execution import CatalogIntegration, CodeExecutionServer
from agora_workbench.data_lake.catalog import CatalogConfig
catalog = CatalogIntegration.development_from_config(
CatalogConfig.from_yaml("catalog.yaml"),
db_path="catalog.db",
)
server = CodeExecutionServer(
server_config=config,
auth_config=auth,
catalog=catalog,
)
The db_path points the server at the SQLite index refreshed in the preceding
CLI steps. development_from_config() is an explicit allow-all shortcut for
local examples. Production servers should use from_config() with an
application authorizer or authorizer_factory.
Here, config and auth are the server configuration and authentication
objects from your existing server. If you do not have one yet, build
your first server before adding the
catalog.
The integration adds these MCP tools:
search_datafinds authorized artifacts.get_artifactreturns metadata for one artifact.list_domainslists visible domain labels.get_catalog_capabilitiesreports the operations available to the session.
Search results may include an opaque load_path. The agent passes that value
unchanged into an execute_*_code call; the session resolves, authorizes, and
caches the bytes before execution. Applications should not parse or construct
load paths themselves.
DevelopmentAllowAllCatalogAuthorizer is only for local development. In
production, provide a caller-aware CatalogAuthorizer or
authorizer_factory, and independently grant the server identity only the
storage permissions it needs.
Move to production deliberately¶
Before deployment:
- Choose
manifestdiscovery when the entire source directory is not approved for disclosure. - Give every source a stable
source_id. - Keep keyword-only search unless vector search is a real requirement.
- Use one rebuildable SQLite index per reader or pod; do not share one writable SQLite file across pods.
- Treat storage RBAC and application authorization as separate controls.
- Use
AuthorizedManagedCatalogWriterfor durable mutations. - Set a refresh process and
max_stale_secondsthat match your freshness requirement. - Monitor source refresh state and catalog readiness.
Where to go next¶
- CLI and quickstarts — local, managed-write, and Azure operating procedures.
- Working with data — compare provisioning, caching, publishing, and catalogs.
- API reference and extension points — implement providers, resolvers, transfer policy, and managed writes.
- Support matrix — supported combinations, limits, and release gates.
- 0.3 migration — upgrade an existing 0.2.x deployment.