Skip to the content.

Data Query

This category evaluates an agent harness and model (or MCP Host) on its ability to retrieve data from a live Business Central environment to answer a natural-language data question. It is execution-based (no LLM judge): the agent reports the rows it retrieved (answer.json), and the run is resolved when those rows match the result set of a hidden gold AL query run against the same Cronus/Contoso demo data (values compared normalized; order ignored unless the entry is ordered).

The point of the category is to compare how the data is retrieved:

Baseline Leaderboard

No results available yet. Check back soon!

BC MCP Experiment

Comparing runs that enable the Business Central Data Query MCP tools (bc-mcp) against the matching no-tooling Default baseline for the same model.

No results available yet. Check back soon!

← Back to Home