Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Importing and Registering Hugging Face Models into Azure ML

This notebook demonstrates the process of importing models from the Hugging Face hub and registering them into Azure Machine Learning (Azure ML) for further use in various machine learning tasks.

Why Hugging Face Models in PyRIT?

The primary goal of PyRIT is to assess the robustness of LLM endpoints against different harm categories such as fabrication/ungrounded content (e.g., hallucination), misuse (e.g., bias), and prohibited content (e.g., harassment). Hugging Face serves as a comprehensive repository of LLMs, capable of generating a diverse and complex prompts when given appropriate system prompt. Models such as:

are particularly useful for generating prompts or scenarios without content moderation. These can be configured as part of a red teaming attack in PyRIT to create challenging and uncensored prompts/scenarios. These prompts are then submitted to the target chat bot, helping assess its ability to handle potentially unsafe, unexpected, or adversarial inputs.

Important Note on Deploying Quantized Models

When deploying quantized models, especially those suffixed with GGML, FP16, or GPTQ, it’s crucial to have GPU support. These models are optimized for performance but require the computational capabilities of GPUs to run. Ensure your deployment environment is equipped with the necessary GPU resources to handle these models.

Supported Tasks

The import process supports a variety of tasks, including but not limited to:

  • Text classification

  • Text generation

  • Question answering

  • Summarization

Import Process

The process involves downloading models from the Hugging Face hub, converting them to MLflow format for compatibility with Azure ML, and then registering them for easy access and deployment.

Prerequisites

  • An Azure account with an active subscription. Create one for free.

  • An Azure ML workspace set up. Learn how to set up a workspace.

  • Install the Azure ML client library for Python with pip.

       pip install azure-ai-ml
       pip install azure-identity
  • Execute the az login command to sign in to your Azure subscription. For detailed instructions, refer to the “Authenticate with Azure Subscription” section here

1. Connect to Azure Machine Learning Workspace

Before we start, we need to connect to our Azure ML workspace. The workspace is the top-level resource for Azure ML, providing a centralized place to work with all the artifacts you create.

Steps:

  1. Import Required Libraries: We’ll start by importing the necessary libraries from the Azure ML SDK.

  2. Set Up Credentials: We’ll use DefaultAzureCredential or InteractiveBrowserCredential for authentication.

  3. Access Workspace and Registry: We’ll obtain handles to our AML workspace and the model registry.

1.1 Import Required Libraries

1.2 Load Environment Variables

Load necessary environment variables from an .env file.

To execute the following job on an Azure ML compute cluster, set AZURE_ML_COMPUTE_TYPE to amlcompute and specify AZURE_ML_INSTANCE_SIZE as STANDARD_D4_V2 (or other as you see fit). When utilizing the model import component, AZURE_ML_REGISTRY_NAME should be set to azureml, and AZURE_ML_MODEL_IMPORT_VERSION can be either latest or a specific version like 0.0.22. For Hugging Face models, the TASK_NAME might be text-generation for text generation models. For default values and further guidance, please see the .env_example file.

Environment Variables

For ex., to download the Hugging Face model cognitivecomputations/Wizard-Vicuna-13B-Uncensored into your Azure environment, below are the environment variables that needs to be set in .env file:

  1. AZURE_SUBSCRIPTION_ID

    • Obtain your Azure Subscription ID, essential for accessing Azure services.

  2. AZURE_RESOURCE_GROUP

    • Identify the Resource Group where your Azure Machine Learning (Azure ML) workspace is located.

  3. AZURE_ML_WORKSPACE_NAME

    • Specify the name of your AZURE ML workspace where the model will be registered.

  4. AZURE_ML_REGISTRY_NAME

    • Choose a name for registering the model in your AZURE ML workspace, such as “HuggingFace”. This helps in identifying if the model already exists in your AZURE ML Hugging Face registry.

  5. HF_MODEL_ID

    • For instance, cognitivecomputations/Wizard-Vicuna-13B-Uncensored as the model ID for the Hugging Face model you wish to download and register.

  6. TASK_NAME

    • Task name for which you’re using the model, for example, text-generation for text generation tasks.

  7. AZURE_ML_COMPUTE_NAME

    • AZURE ML Compute where this script runs, specifically an Azure ML compute cluster suitable for these tasks.

  8. AZURE_ML_INSTANCE_SIZE

    • Select the size of the compute instance of Azure ML compute cluster, ensuring it’s at least double the size of the model to accommodate it effectively.

  9. AZURE_ML_COMPUTE_NAME

    • If you already have an Azure ML compute cluster, provide its name. If not, the script will create one based on the instance size and the specified minimum and maximum instances.
      AML compute cluster

  10. IDLE_TIME_BEFORE_SCALE_DOWN

    • Set the duration for the Azure ML cluster to remain active before scaling down due to inactivity, ensuring efficient resource use. Typically, 3-4 hours is ideal for large size models.

1.3 Configure Credentials

Set up the DefaultAzureCredential for seamless authentication with Azure services. This method should handle most authentication scenarios. If you encounter issues, refer to the Azure Identity documentation for alternative credentials.

1.4 Access Azure ML Workspace and Registry

Using the Azure ML SDK, we’ll connect to our workspace. This requires having a configuration file or setting up the workspace parameters directly in the code. Ensure your workspace is configured with a compute instance or cluster for running the jobs.

1.5 Compute Target Setup

For model operations, we need a compute target. Here, we’ll either attach an existing AmlCompute or create a new one. Note that creating a new AmlCompute can take approximately 5 minutes.

  • Existing AmlCompute: If an AmlCompute with the specified name exists, we’ll use it.

  • New AmlCompute: If it doesn’t exist, we’ll create a new one. Be aware of the resource limits in Azure ML.

Important Note for Azure ML Compute Setup:

When configuring the Azure ML compute cluster for running pipelines, please ensure the following:

  1. Idle Time to Scale Down: If there is an existing Azure ML compute cluster you wish to use, set the idle time to scale down to at least 4 hours. This helps in managing compute resources efficiently to run long-running jobs, helpful if the Hugging Face model is large in size.

  2. Memory Requirements for Hugging Face Models: When planning to download and register a Hugging Face model, ensure the compute size memory is at least double the size of the Hugging Face model. For example, if the Hugging Face model size is around 32 GB, the Azure ML cluster node size should be at least 64 GB to avoid any issues during the download and registration process.

2. Create an Azure ML Pipeline for Hugging Face Models

In this section, we’ll set up a pipeline to import and register Hugging Face models into Azure ML.

Steps:

  1. Load Pipeline Component: We’ll load the necessary pipeline component from the Azure ML registry.

  2. Define Pipeline Parameters: We’ll specify parameters such as the Hugging Face model ID and compute target.

  3. Create Pipeline: Using the loaded component and parameters, we’ll define the pipeline.

  4. Execute Pipeline: We’ll submit the pipeline job to Azure ML and monitor its progress.

2.1 Load Pipeline Component

Load the import_model pipeline component from the Azure ML registry. This component is responsible for downloading the Hugging Face model, converting it to MLflow format, and registering it in Azure ML.

2.2 Create and Configure the Pipeline

Define the pipeline using the import_model component and the specified parameters. We’ll also set up the User Identity Configuration for the pipeline, allowing individual components to access identity credentials if required.

2.3 Submit the Pipeline Job

Submit the pipeline job to Azure ML for execution. The job will import the specified Hugging Face model and register it in Azure ML. We’ll monitor the job’s progress and output.