Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Deploying Hugging Face Models into Azure ML Managed Online Endpoint

This notebook demonstrates the process of deploying registered models in Azure ML workspace to an AZURE ML managed online endpoint for real-time inference.

Learn more about Azure ML Managed Online Endpoints

Prerequisites

  • An Azure account with an active subscription. Create one for free.

  • An Azure ML workspace set up. Learn how to set up a workspace.

  • Install the Azure ML client library for Python with pip.

       pip install azure-ai-ml
       pip install azure-identity
  • Execute the az login command to sign in to your Azure subscription. For detailed instructions, refer to the “Authenticate with Azure Subscription” section in the markdown file provided here

  • A Hugging Face model should be present in the AZURE ML model catalog. If it is missing, execute the notebook to download and register the Hugging Face model in the AZURE ML registry.

Load Environment Variables

Load necessary environment variables from an .env file.

For example, to download the Hugging Face model cognitivecomputations/Wizard-Vicuna-13B-Uncensored into your Azure environment, below are the environment variables that needs to be set in .env file:

  1. AZURE_SUBSCRIPTION_ID

    • Obtain your Azure Subscription ID, essential for accessing Azure services.

  2. AZURE_RESOURCE_GROUP

    • Identify the Resource Group where your Azure Machine Learning (AZURE ML) workspace is located.

  3. AZURE_ML_WORKSPACE_NAME

    • Specify the name of your AZURE ML workspace where the model will be registered.

  4. AZURE_ML_REGISTRY_NAME

    • Choose a name for registering the model in your AZURE ML workspace, such as “HuggingFace”. This helps in identifying if the model already exists in your AZURE ML Hugging Face registry.

  5. AZURE_ML_MODEL_NAME_TO_DEPLOY

    • If the model is listed in the AZURE ML Hugging Face model catalog, then supply the model name as shown in the following image.
      AML Hugging Face model

    • If you intend to deploy the model from the AZURE ML workspace model registry, then use the model name as shown in the subsequent image.
      AML Workspace model

  6. AZURE_ML_MODEL_VERSION_TO_DEPLOY

    • You can find the details of the model version in the images from previous step associated with the respective model.

  7. AZURE_ML_MODEL_DEPLOY_INSTANCE_SIZE

    • Select the size of the compute instance of for deploying the model, ensuring it’s at least double the size of the model to effective inference.

  8. AZURE_ML_MODEL_DEPLOY_INSTANCE_COUNT

    • Number of compute instances for model deployment.

  9. AZURE_ML_MODEL_DEPLOY_REQUEST_TIMEOUT_MS

    • Set the AZURE ML inference endpoint request timeout, recommended value is 60000 (in millis).

  10. AZURE_ML_MODEL_DEPLOY_LIVENESS_PROBE_INIT_DELAY_SECS

    • Configure the liveness probe initial delay value for the Azure ML container hosting your model. The default initial_delay value for the liveness probe, as established by Azure ML managed compute, is 600 seconds. Consider raising this value for the deployment of larger models.

Configure Credentials

Set up the DefaultAzureCredential for seamless authentication with Azure services. This method should handle most authentication scenarios. If you encounter issues, refer to the Azure Identity documentation for alternative credentials.

Create an Azure ML managed online endpoint To define an endpoint, you need to specify:

Endpoint name: The name of the endpoint. It must be unique in the Azure region. For more information on the naming rules, see managed online endpoint limits. Authentication mode: The authentication method for the endpoint. Choose between key-based authentication and Azure Machine Learning token-based authentication. A key doesn’t expire, but a token does expire.

Add deployment to an Azure ML endpoint created above

Please be aware that deploying, particularly larger models, may take some time. Once the deployment is finished, the provisioning state will be marked as ‘Succeeded’, as illustrated in the image below.
AML Endpoint Deployment