Skip to main content

Prerequisites and Build Validation

[!NOTE] This guide expands on the Prerequisites section of the main contributing guide.

Tools, Azure access, and build validation requirements for contributing to the Physical AI Toolchain.

Required Tools

Install these tools before contributing:

ToolMinimum VersionInstallation
Terraform1.9.8https://developer.hashicorp.com/terraform/install
TFLint0.61.0https://github.com/terraform-linters/tflint
Azure CLI2.65.0https://learn.microsoft.com/cli/azure/install-azure-cli
kubectl1.31https://kubernetes.io/docs/tasks/tools/
Helm4.0https://helm.sh/docs/intro/install/
Node.js/npm20+ LTShttps://nodejs.org/
Python3.12+https://www.python.org/downloads/
shellcheck0.10+https://www.shellcheck.net/
uvlatesthttps://docs.astral.sh/uv/
Go1.24+https://go.dev/dl/
golangci-lint2.11+https://golangci-lint.run/welcome/install/
Dockerlatesthttps://docs.docker.com/get-docker/ (with NVIDIA Container Toolkit)
OSMO CLIlatesthttps://developer.nvidia.com/osmo
terraform-docs0.21.0https://github.com/terraform-docs/terraform-docs/releases
OSV-Scanner2.3.8https://github.com/google/osv-scanner/releases/tag/v2.3.8 (installed automatically by setup-dev.sh / setup-dev.ps1)
hve-corelatesthttps://github.com/microsoft/hve-core

[!NOTE] GitHub Copilot Coding Agent runs in a separate cloud GitHub Actions environment provisioned by .github/workflows/copilot-setup-steps.yml. When you bump a language runtime or test runner version locally (devcontainer or this list), update the matching pin in that workflow so cloud-agent sessions stay aligned.

Dev Container GPU Runtime

The dev container requests a GPU only when Dev Containers detects one. Configure the host according to its platform:

HostDocker runtime configuration
x86_64 with an NVIDIA GPUInstall NVIDIA Container Toolkit; automatic --gpus all handling normally requires no change
NVIDIA ARM64 host in CSV modeSet nvidia as Docker's default runtime
Host without an NVIDIA GPURetain the default runc runtime; the dev container starts without GPU devices

On NVIDIA ARM64 hosts, the NVIDIA Container Runtime can auto-detect CSV mode and reject the --gpus all request unless Docker invokes the nvidia runtime. Confirm that Docker has registered the runtime:

docker info --format '{{json .Runtimes}}'

The output must include nvidia. Configure it as the default runtime, then reload Docker without stopping running containers:

sudo nvidia-ctk runtime configure --runtime=docker --set-as-default
sudo systemctl reload docker

If the reload does not apply the change, restart Docker. A restart interrupts running containers:

sudo systemctl restart docker

Verify the default runtime and GPU attachment:

docker info --format 'Default runtime: {{.DefaultRuntime}}'
docker run --rm --runtime=nvidia --gpus all ubuntu:24.04 \
sh -c 'test -e /dev/nvidia0 && echo "GPU attached"'

The first command must report nvidia, and the second must print GPU attached. Rebuild and reopen the dev container after configuring the host.

This configuration resolves the following container startup error:

invoking the NVIDIA Container Runtime Hook directly (e.g. specifying the docker --gpus flag) is not supported.
Please use the NVIDIA Container Runtime (e.g. specify the --runtime=nvidia flag) instead

GPU Device Group Access

Attaching the GPU is not sufficient. CUDA also requires the container user to hold the host groups that own the GPU device nodes. .devcontainer/devcontainer.json grants them through runArgs:

Device nodeHost groupRequired for
/dev/nvmapvideoTegra memory manager on NVIDIA ARM64 hosts
/dev/dri/renderD*renderDRM render node that CUDA opens

The committed configuration assumes the host render group is GID 993. Confirm the value on the host:

stat -c '%g %G' /dev/dri/renderD128

If the GID differs, update the matching --group-add entry and rebuild the dev container. Group names do not transfer across the container boundary, so the numeric GID is required; inside the container it resolves to whichever name holds that GID, commonly systemd-resolve.

Missing membership produces a misleading failure. nvidia-smi lists the GPU and torch.cuda.device_count() returns 1, but torch.cuda.is_available() returns False:

CUDA initialization: Unexpected error from cudaGetDeviceCount(). Error 801: operation not supported

Verify device access and CUDA inside the dev container:

test -r /dev/nvmap && test -r /dev/dri/renderD128 && echo "device access OK"
python -c "import torch; print('cuda:', torch.cuda.is_available())"

Both commands must succeed, and the second must print cuda: True.

Azure Access Requirements

Deploying this architecture requires Azure subscription access with specific permissions and quotas:

Subscription Roles

  • Contributor role for resource group creation and management
  • User Access Administrator role for managed identity assignment

GPU Quota

  • Request GPU VM quota in your target region before deployment
  • Architecture uses Standard_NC24ads_A100_v4 (24 vCPU, 220 GB RAM, 1x A100 80GB GPU)
  • Check quota: az vm list-usage --location <region> --query "[?name.value=='standardNCadsA100v4Family']"
  • Request increase through Azure Portal → Quotas → Compute

Regional Availability

NVIDIA NGC Account

Training workflows use NVIDIA GPU Operator and Isaac Lab, which require NGC credentials:

  • Create account: https://ngc.nvidia.com/signup
  • Generate API key: NGC Console → Account Settings → Generate API Key
  • Store API key in Azure Key Vault or Kubernetes secret (deployment scripts provide guidance)

Cost Awareness

Full deployment validation incurs Azure costs. Understand cost structure before deploying:

GPU Virtual Machines

  • Standard_NC24ads_A100_v4: ~$3.06/hour per VM (pay-as-you-go)
  • 8-hour validation session: ~$25
  • 40-hour work week: ~$125

Managed Services

  • AKS control plane: $0.10/hour ($73/month)
  • Log Analytics workspace: ~$2.76/GB ingested
  • Storage accounts: ~$0.02/GB (block blob, hot tier)
  • Azure Container Registry: Basic tier ~$5/month

Cost Optimization

  • Use terraform destroy immediately after validation
  • Automate cleanup with -auto-approve flag
  • Monitor costs: Azure Portal → Cost Management + Billing
  • Set budget alerts to prevent overruns

Estimated Costs

  • Quick validation (deploy + verify + destroy): ~$25-50
  • Extended development session (8 hours): ~$50-100
  • Monthly development (40 hours): ~$200-300

Build and Validation Requirements

Tool Version Verification

Verify tool versions before validating:

# Terraform
terraform version # >= 1.9.8

# TFLint (Terraform linter)
tflint --version # >= 0.61.0

# Azure CLI
az version # >= 2.65.0

# kubectl
kubectl version --client # >= 1.31

# Helm
helm version # >= 3.16

# Node.js (for documentation linting)
node --version # >= 20

# Python (for training scripts)
python --version # >= 3.12

# shellcheck (for shell script validation)
shellcheck --version # >= 0.10

# uv (Python package manager)
uv --version

# Go
go version # >= 1.24

# golangci-lint
golangci-lint version # >= 2.11

# Docker with NVIDIA Container Toolkit
docker --version
nvidia-ctk --version

# OSMO CLI
osmo --version

# terraform-docs
terraform-docs --version # >= 0.21.0

# OSV-Scanner (dependency vulnerability scanner)
osv-scanner --version # == 2.3.8 (pinned; installed by setup-dev scripts)

# hve-core (VS Code extension — verify via extensions list)
code --list-extensions | grep -i hve-core

TFLint Local Setup

Install TFLint v0.61.0 or newer before changing Terraform modules:

# macOS
brew install tflint

# Linux
curl -s https://raw.githubusercontent.com/terraform-linters/tflint/master/install_linux.sh | bash
# Windows (Chocolatey)
choco install tflint

# Windows (Scoop)
scoop install tflint

Initialize the repository TFLint plugins once from the repository root. This downloads the Azure provider ruleset declared in .tflint.hcl:

tflint --init

Then run the project wrapper before pushing Terraform changes:

npm run lint:tf

The wrapper runs TFLint recursively against infrastructure/terraform/ with the shared .tflint.hcl configuration. A VS Code TFLint extension is optional for inline diagnostics, but the CLI setup above remains the required validation path.

Validation Commands

Run these commands before committing:

Terraform:

# Format check (required)
terraform fmt -check -recursive infrastructure/terraform/

# Initialize and validate (required for infrastructure changes)
cd infrastructure/terraform/
terraform init
terraform validate

# Lint Terraform configurations (required for infrastructure changes)
tflint --init # first time only, installs plugins from .tflint.hcl
tflint --recursive infrastructure/terraform/

Shell Scripts:

# Lint all shell scripts (required)
shellcheck deploy/**/*.sh scripts/**/*.sh

Go:

# Lint Go modules (required for Go changes)
npm run lint:go

# Test Go modules (required for Go changes)
npm run test:go

# Contract tests (validates Terraform outputs against Go struct — requires terraform-docs)
# Run after adding/removing/renaming Terraform outputs
./infrastructure/terraform/e2e/run-contract-tests.sh

Documentation:

# Install dependencies (first time only)
npm install

# Lint markdown (required for documentation changes)
npm run lint:md

VS Code Configuration

The workspace is configured with python.analysis.extraPaths pointing to src/, enabling imports like:

from training.utils import AzureMLContext, bootstrap_azure_ml

Select the .venv/bin/python interpreter in VS Code for IntelliSense support.

The workspace .vscode/settings.json also configures Copilot Chat to load instructions, prompts, and chat modes from hve-core:

Settinghve-core Paths
chat.modeFilesLocations../hve-core/.github/chatmodes, ../hve-core/copilot/beads/chatmodes
chat.instructionsFilesLocations../hve-core/.github/instructions, ../hve-core/copilot/beads/instructions
chat.promptFilesLocations../hve-core/.github/prompts, ../hve-core/copilot/beads/prompts

These paths resolve when hve-core is installed as a peer directory or via the VS Code Extension. Without hve-core, Copilot still functions but shared conventions, prompts, and chat modes are unavailable.

For a complete list of available agents, prompts, and skills, see Copilot Artifacts.