Development Guide
Start with CONTRIBUTING.md for issue-first scope approval, PR eligibility, small corrections, private security reporting, and the transition for existing contributions. GOVERNANCE.md names the responsible human maintainers and explains decisions and responsibilities. Its roadmap and release planning model separates priorities from scope approval and release targets. Follow the public Roadmap for current Now / Next / Later priorities. Those root documents own contribution policy; this guide covers technical work within the approved scope. Reporting and investigation need no prior permission.
Development Environment
Section titled “Development Environment”This project uses uv to manage Python environments and dependencies:
# Clone the repositorygit clone https://github.com/microsoft/apm.gitcd apm
# Install all dependencies (creates .venv automatically)uv sync --extra devFor binary builds, install the locked build tools with
uv sync --frozen --extra dev --extra build. PyInstaller 6.17 or newer is
required for compatibility with setuptools 83, which no longer provides
pkg_resources.
Optional agent tools
Section titled “Optional agent tools”No AI tool, harness, or repository skill is required to contribute. APM
dogfoods its own primitives from
.apm/skills/ and
.apm/agents/, with
local package dependencies declared in
apm.yml.
Automated recommendations are advisory, not scope approval or a substitute
for human review.
Issue triage produces a recommendation and a proposed scope, done-when,
exclusions, and review-needs brief. Maintainers still decide acceptance,
priority, contributor invitations, and milestones. The
triage label contract
separates those decisions from advisory processing. During compatibility
rollout, both status/triaged and triage/recommended mean completed
automated advice, not human review. Sweep fetch excludes those labels
so already-advised open issues and PRs are not re-listed. The new writer
uses triage/recommended after the canonical label is provisioned and
this code is deployed. Do not write status/triaged.
status/needs-triage can remain after advice while awaiting a human decision.
No label or milestone migration is performed by the advisory workflow.
PR review is the same shape: panel-review requests a fresh
advisory pass (autopilot-pr-review-worker). status/accepted
(on the PR or a linked issue) is the human action flag – no
accepted, no review. If acceptance is missing, scheduler and
review-worker stop with no comment. The worker may clear
panel-review; the scheduler does not comment or change labels.
The review is advisory and does not gate merge. Drive-to-merge
(autopilot-pr-merge-worker) is summoned by name and is never
composed by the review scheduler.
Eligibility evidence and automation
Section titled “Eligibility evidence and automation”The scope record format
provides explicit human evidence on an issue. The deterministic
PR eligibility (advisory only) check is always neutral: record-present,
withdrawn, needs-evidence, or error are not implementation permission.
Private security/dependency tracking and claimed trivial corrections receive
manual-review states. Real issue references are checked through GitHub, not
inferred from arbitrary numbers, PR links, upstream references, or labels.
Humans still decide whether the change matches the linked scope.
After the workflow reaches the default branch, PR updates and relevant
issue/comment changes refresh evidence without invoking an LLM panel.
Forks and stacked PRs use default-branch governance and code, never the head
or a feature-branch approval roster. A permissionless queue signal wakes the
default-branch reporter through workflow_run; the reporter verifies GitHub’s
run, repository, event, and workflow metadata without consuming artifacts or
contributor code. Commit associations do not prove complete queue membership;
missing association/read access is reported as unknown/error. The check is
not part of the required merge gate, and this PR does not deploy itself.
Native event delivery and queue limits still apply. Maintainers can request a
default-branch recheck without selecting a branch-controlled workflow definition:
gh api --method POST repos/microsoft/apm/dispatches \ -f event_type=pr-eligibility-recheck -F 'client_payload[pr_number]=123'There is no workflow_dispatch entrypoint. These workflow-definition
guarantees target GitHub.com, not older GitHub Enterprise Server versions.
For a read-only local assessment, run the tool from a trusted default-branch
checkout with Node.js and an authenticated gh installed outside the project.
The tool excludes project PATH entries and symlinks back into the project:
node scripts/governance/eligibility.cjs --repo microsoft/apm --pr 123node scripts/governance/eligibility.cjs --helpThe tool reads current policy from the default branch. Until the new roster format is deployed there, it deliberately reports a policy error rather than using a feature branch as authority. A snapshot cannot recover deleted withdrawals, so manual automation still requires fresh responsible-human confirmation of the bounded issue scope. No historical acceptance is restored.
Each assessment is bounded: 25 visible unique issue references per PR, 10
metadata pages per collection, 100 matched PR targets, 500 total metadata reads,
and 32 MiB of metadata per operation. Issue events may inspect more than 100
open PRs; only matched targets count toward that cap. A full final page is incomplete evidence.
Exceeding a bound reports error, never a truncated record-present result.
Reads are cached only within one operation; the final PR head/body refresh
bypasses that cache. This accommodates the current 62-PR policy refresh while
bounding genuinely oversized operations. Hidden HTML comments are not issue
traceability or approval nominations.
Once deployed, Daily Docs Updater runs discovery only, without issue/PR creation or auto-merge capabilities. Scheduled and manual runs stop at a human handoff. A docs-sync confirmation label likewise requests consideration; it cannot authorize a companion PR. Bug sweeps read canonical and legacy bug labels once per issue, then require the same human checkpoint before implementation. During rollout, keep Daily Docs disabled until the replacement source and compiled capabilities are deployed and verified. Opening or merging the PR does not authorize resuming the disabled workflow; a maintainer does that separately.
The eligibility workflow verifies that main is still the repository default
branch before checking out that literal ref, then uses the checked-out commit
for both implementation and policy. A default-branch rename stops the workflow
until maintainers review this boundary; it never falls back to a PR or event ref.
Installing optional skills
Section titled “Installing optional skills”To use these tools, install APM if needed, then run from the repository root:
apm installThe manifest’s includes: auto picks up .apm/. Its pinned copilot target
deploys to the committed .github/ and .agents/skills/ tree, regardless of
which harness your machine detects. Your harness can then discover and invoke
the installed skills by name.
For a different harness, exclude its generated roots locally before overriding the pinned target. For example, with Claude Code:
printf '.claude/\n' >> "$(git rev-parse --git-path info/exclude)"apm install --target claudeThe local exclusion avoids changing repository-wide ignore rules. This install
also adds deploy paths to apm.lock.yaml; leave that local override uncommitted.
Check the target catalogue
for other targets; some write to more than one root.
Testing
Section titled “Testing”After setup, use pytest for focused feedback:
# Unit suiteuv run pytest tests/unit tests/test_console.py -x
# Focused file; replace with the test relevant to your changeuv run pytest tests/test_console.py -x
# Full suite, including integration and acceptance testsuv run pytest
# Verbose unit resultsuv run pytest tests/unit -x -vpytest-xdist is available: add -n auto for parallel execution or -n0
to force serial execution. The default selection in pyproject.toml excludes
benchmark and live tests.
Without uv, use a standard Python venv and pip:
# create and activate a venv (POSIX / WSL)python -m venv .venvsource .venv/bin/activate
# install this package in editable mode and test depspip install -U pippip install -e '.[dev]'
# run unit testspytest tests/unit tests/test_console.py -xRunning integration tests
Section titled “Running integration tests”Tests under tests/integration/ declare preconditions with requires_*
markers. The _MARKER_CHECKS registry in tests/integration/conftest.py
skips tests with missing prerequisites at collection time and reports why.
Use the marker registry to
find the token, runtime setup command, or opt-in flag each family needs.
# Run tests whose prerequisites your environment satisfiesuv run pytest tests/integration -v
# Select a prerequisite familyuv run pytest tests/integration -m requires_github_token -vWhen adding a precondition, add its check to _MARKER_CHECKS and declare the
marker in pyproject.toml; do not duplicate the check in each test. For install,
compile, pack, or audit lifecycle changes, reuse the
hermetic lifecycle fixtures
to exercise the real CLI with a sanitized child environment and reviewed local
Git sources.
Coverage policy
Section titled “Coverage policy”Both suites have hard CI coverage gates that must pass before merge:
| Suite | Gate | Enforced in |
|---|---|---|
| Unit | 80% | pyproject.toml (fail_under) and combined coverage in .github/workflows/ci.yml |
| Integration | 70% | Combined coverage in .github/workflows/ci-integration.yml (--fail-under) |
Gates only move upward. When actual coverage exceeds the gate by at least
5 percentage points, raise the gate to actual - 3 in the next release PR.
CI’s coverage summaries include a “Lowest-coverage files” section, rendered by
scripts/coverage-summary.py, to identify where new tests would help.
Running the bounded mutation pilot
Section titled “Running the bounded mutation pilot”The advisory mutation pilot covers five stable owners: dependency subset
selection, update-plan construction, cached-policy serialization, canonical
in-package link projection, and lockfile field normalization (the fail-closed
host_type/exec_status normalizers, not the @dataclass reconstruction
methods to_dict/from_dict/to_dependency_ref – mutmut cannot mutate
@dataclass methods; those are defended by PR #2246’s manual mutation-break
twins instead). It runs nightly or by manual workflow
dispatch, not as required PR CI, and has a 20-minute hosted job budget.
Run the exact-function allowlist locally:
uv run --frozen --extra dev python scripts/run_mutation_pilot.py \ --output mutation-pilot-report.jsonThe command fails on new survivors, timeouts, suspicious results, unchecked
mutants, and incomplete outcomes. It writes a sorted, timestamp-free JSON
report even when survivor comparison fails. Pass --reuse-cache only when
the allowlisted source, test seams, configuration, runner, and lockfile
are unchanged.
To inspect existing mutmut metadata without executing mutants:
uv run --frozen --extra dev python scripts/run_mutation_pilot.py \ --report-only --output mutation-pilot-report.jsonThe reviewed survivor allowlist lives in
tests/mutation/baseline.json.
Do not update it to make a run green. Review every surviving diff with
mutmut show and add behavioral tests for real contract gaps.
Use --update-baseline only when the baseline change itself has been reviewed:
uv run --frozen --extra dev python scripts/run_mutation_pilot.py \ --update-baseline --output mutation-pilot-report.jsonCoding Style
Section titled “Coding Style”APM follows PEP 8 and uses
Ruff for linting and formatting. Follow the
canonical lint contract
for local commands, auto-fixes, and common diagnostics. The actual
Lint job in ci.yml
defines the complete enforced step list; mirror it before pushing or claiming
green CI.
Ruff lint and format are only part of that job. It also checks YAML I/O,
file length, portable relative paths, duplication, auth-protocol boundaries,
and architecture boundaries. Its Python scope includes the architecture-linter
scripts as well as src/ and tests/. CI checks the PR merge result, so changes
on main can introduce failures even when the branch alone passes.
Architecture guardrails
Section titled “Architecture guardrails”Durable architecture decisions have one canonical owner. Executable owner metadata
lives in .apm/architecture/owners/index.json and the six shards it lists:
core-runtime.jsoninstall-deployment.jsonhooks-integrations.jsontransport-auth-platform.jsonmarketplace-plugins.jsoncontracts-tooling.json
For an ordinary owner addition, edit the appropriate shard. Do not edit
.apm/instructions/architecture.instructions.md, .github/instructions/...
files, or apm.lock.yaml. Each owner entry has id, decision, owner,
selectors, and guards fields; this metadata is an ownership registry, not a
rule DSL.
When centralizing behavior, add a behavioral test and register a semantic static guard. Run the stable architecture check with:
bash scripts/lint-architecture-boundaries.shThe check fails closed when metadata is malformed, missing, or not listed in the index.
InstallTransaction owns one acquisition of the shared lifecycle FileLock.
Commit, rollback, and context exit are the normal release paths; explicit release
invokes the same weakref.finalize callback used for abandoned transactions.
This fallback releases only the transaction’s outstanding acquisition, never
runs filesystem rollback, and leaves other owners’ acquisitions intact. Keep
transaction lifetime and completion on the acquiring thread: filelock uses
thread-local state, so there is no cross-thread lifecycle guarantee. In lifecycle
release regression tests, retain the shared lock handle so FileLock destruction
cannot mask a missed release.
Optional: local pre-commit hooks
Section titled “Optional: local pre-commit hooks”For instant feedback before pushing, install the pre-commit hooks:
uv run pre-commit installThis is optional – CI is the authoritative gate. The pre-commit hook rev may lag behind the CI version; check .pre-commit-config.yaml against uv.lock if you see discrepancies.
CI and merging
Section titled “CI and merging”How merging works
Section titled “How merging works”A maintainer adds an approved PR to GitHub’s native merge queue. The queue
builds a tentative merge against the latest main, runs checks including the
integration suite, and merges on success or ejects the PR on failure.
There is no manual “Update branch” step just to enter the queue. If a real
failure ejects your PR, push a fix and ask a maintainer to re-queue it.
Fast unit and build checks (Tier 1) run on PR updates. The required Lifecycle
Smoke check also runs on PRs and merge-queue commits. It selects
lifecycle_smoke and not lifecycle_merge_group contracts with no network,
credentials, or frozen binary required. See
Integration Testing for the bounded selection,
timeout, prerequisites, and local command.
The full integration suite (Tier 2) runs in the queue rather than on every
WIP push.
Workflow dependency updates
Section titled “Workflow dependency updates”When updating actions in generated .github/workflows/*.lock.yml files,
align gh-aw-manifest headers, action lists, and .github/aw/actions-lock.json
entries with runtime uses: pins. Dependabot does not update this metadata.
Preserve compiler versions and source hashes for dependency-only edits.
For source changes, including comments, run gh aw compile with the repository’s
pinned compiler version from .github/workflows/copilot-setup-steps.yml.
Commit source and regenerated lock together; do not edit compiler metadata by hand.
Run uv run --frozen --extra dev pytest tests/unit/test_triage_panel_lock.py tests/unit/test_shared_apm_workflow_contract.py
to check triage/review source freshness across line endings, compiler-header
consistency with the repository pin, and action-pin consistency.
Code scanning on pull requests and merge queues
Section titled “Code scanning on pull requests and merge queues”The CodeQL workflow runs Python and GitHub Actions analysis on pull requests,
pushes to main, merge-queue checks_requested events, and the weekly schedule.
Keep the workflow path, analyze job ID, and language matrix stable: they
identify the analysis configurations GitHub compares against the base branch.
PR results do not replace results for the merge queue’s separate commit.
If both analysis jobs succeed but Code scanning still reports a missing
configuration, inspect the CodeQL check summary. An additional API upload
configuration on the base branch belongs to a separate upload producer;
rerunning this workflow cannot supply that producer’s results. Coordinate
matching PR and queue uploads with its owner rather than deleting findings,
renaming categories, or weakening the code-scanning ruleset.
Documentation
Section titled “Documentation”If your changes affect how users interact with the project, update the documentation accordingly.
Public top-level CLI commands and rendered reference pages are a matched contract. When you add,
remove, or rename a command, create, remove, or rename its matching page under
docs/src/content/docs/reference/cli/ and update the command table in
docs/src/content/docs/reference/index.md, then run:
npm --prefix docs run builduv run --frozen python scripts/check_cli_docs.py docs/distExtending APM
Section titled “Extending APM”Adding or modifying an MCP client adapter
Section titled “Adding or modifying an MCP client adapter”Adapters in src/apm_cli/adapters/client/ inherit shared utilities from
MCPClientAdapter in base.py:
- Reuse
_apply_pypi_homebrew_generic_config,_apply_auth_and_headers_impl, and_resolve_env_vars_with_promptingrather than copying sibling adapters. - The pylint R0801 similarity threshold is 10 lines; duplicated blocks fail CI.
- For marketplace tag parsing, use
marketplace._shared.iter_semver_tagsrather than reimplementing the refs-iteration loop.
How to add an experimental feature flag
Section titled “How to add an experimental feature flag”Use an experimental flag for a user-visible behavior change that needs early adopter feedback, not a bug fix, internal refactor, or change that should ship as the default. Flags are ergonomic/UX toggles only. They MUST NOT gate security-critical behavior: content scanning, path validation, lockfile integrity, token handling, MCP trust, or collision detection.
-
Register the flag in
src/apm_cli/core/experimental.py’sFLAGSdict with a frozenExperimentalFlag(name=..., description=..., default=False, hint=...). -
Import and call
is_enabledat function scope to avoid import cycles and config I/O at module import time. For example, the existingverbose_versionflag can be checked with:def show_runtime_details():from apm_cli.core.experimental import is_enabledreturn is_enabled("verbose_version") -
Test both enabled and disabled paths.
-
Update the experimental command reference.
Use snake_case in the registry and config, and kebab-case for display and
other user-facing strings. The CLI accepts both forms on input. Persist flag
state only in ~/.apm/config.json through update_config.
When a flag graduates to the default, remove its gate and FLAGS entry in the
same PR. Add a CHANGELOG.md entry under Changed, with a migration note if
the previous default differed.
Adding or changing a normative requirement (OpenAPM v0.1)
Section titled “Adding or changing a normative requirement (OpenAPM v0.1)”The OpenAPM v0.1 specification and APM’s implementation evolve together. Every normative change MUST include three coupled edits in the same PR:
- Spec: add or change a
<a id="req-XXX"></a>anchor and its prose indocs/src/content/docs/specs/openapm-v0.1.md, plus the matching Appendix C row. - Manifest: update
docs/src/content/docs/specs/manifests/openapm-v0.1.requirements.ymlto remain a byte-equivalent projection of the canonical anchors. - Test: add or extend a
@pytest.mark.req("req-XXX")test undertests/spec_conformance/. If a real assertion is not yet possible, callwaive("...")from_helpers.pywith a one-line rationale. The waiver appears inCONFORMANCE.mdas visible debt, not test coverage.
Regenerate the conformance statement after these edits:
uv run --extra dev python -m tests.spec_conformance.gen_statementInclude the resulting root CONFORMANCE.md and CONFORMANCE.json in the
same PR; CI requires a clean generated diff.
The workflow distinguishes three modes:
- Mode A (silent regression): a code change breaks an assertion bound to
a
req-XXX. The spec-conformance pytest job fails. Fix the code, not the spec. - Mode B (silent extension): new behavior under a normative critical path
lacks a spec citation. The four-way
orphan_checkcatches a requirement marker missing its anchor, manifest row, or Appendix C row. The Mode B detector catches substantive critical-path code with no spec artifacts at all. Add the anchor, manifest row, and marker, with the Appendix C row as above. For a true refactor, performance rewrite, or internal cleanup with no observable behavior change, addapm-spec-waiver: <one-line rationale>to the PR body or a commit message. The rationale must be at least 16 characters; CI echoes the waiver verbatim for reviewer inspection. The critical-path allowlist lives intests/spec_conformance/critical_paths.txt; changes to that list are themselves critical-path edits. - Mode C (stale spec): the prose misstates intended behavior. Amend the anchor, Appendix C row, and manifest entry, with a test proving the intended behavior in the same PR.
These checks cannot detect every semantic drift. Choosing the appropriate mode remains a human decision; the harness exposes the choice, not the answer.