Hands On Lab
MODULE I: DATA INGESTION
COVID-19 Dataset
MODULE 1:
Data Ingestion with Azure Data Factory
Pre-requisites:
-
Azure subscription with Azure Data Factory Instance
-
Completed Objective 1 of Module 1: Creating a pipeline to copy country code dataset to the Azure Data Lake
Learning Outcomes for Module 1:
-
Importing COVID-19 Data from Github
-
Creating a linked service to link your Azure Storage account to the data factory. The linked service has the connection information that the Data Factory service uses at runtime to connect to it
-
Creating and validating a pipeline with specified parameters to copy the COVID-19 data to Azure Data Lake
- Create a new pipeline through selecting the plus (+) button and click on Pipeline
- In the general tab, specify the pipeline name i.e. Covid Latest Pipe. You will see the name of your pipeline appear on the left hand panel under Factory Resources.
- We are working with the daily reports dataset in this section, in which a new file is created everyday. We must therefore define a parameter for our pipeline pertaining to the file date of the particular file we want to work with. Navigate to the Parameters tab and click New.
- Input FileDate as the parameter Name. This is the parameter we will use for the rest of this section.
- Next, click on the move and transform section and drag the ‘copy data’ function and rename it to be ‘CovidDataCopy’ in the General panel.
- The dataset we will be working with for this module can be found here: https://github.com/CSSEGISandData/COVID-19/tree/master/csse_covid_19_data/csse_covid_19_daily_reports
- Next up, we are going to create a new source. Click on the source tab and select the (+) button to add a new dataset.
- On the New Dataset panel on the right hand side, select ‘HTTP’ and click continue.
- Given our raw data from Github is in a csv file click the DelimitedText option and click Continue.
- In the Set Properties panel rename the file to ‘CovidDataSource’.
-
Under the Linked Service heading click ‘new’.
-
Name the linked service to ‘HTTPServerReferenceLinkedService’
-
Copy the base URL of the raw Github csv file as shown below https://raw.githubusercontent.com/CSSEGISandData/COVID-19/master/csse_covid_19_data/csse_covid_19_daily_reports/
-
Set authentication to Anonymous
- Click Test connection to verify that the connection is successful.
-
Click Create.
-
Click OK.
-
Click Open Source dataset
- Navigate to the Connection tab. We will be adding dynamic content under the Relative URL field to refer to the file source.
Before we can run the file source however, another parameter must be created to contain extra bits of information. Navigate to the Parameters tab, and input FileName in the Name field.
- Navigate back to the Connection tab and click Add dynamic content under the Relative URL field. What we want to do here is get the FileDate as a csv. Input the following in the add dynamic content field:
We essentially want this dynamic content to be referred to every time the pipe is run.
-
Click Finish
-
Return back to your actual pipeline to test that our parameters work. Under the fileName field, we must input a value to refer to a specific file. This file name is going to come from the pipeline, here we will be using the FileDate parameter we specified at the beginning of this section. Click on Add dynamic content under the fileName field and input the following:
-
Click Finish
-
Next, click on Preview data and input a value for the file date as follows:
Through this process we have learnt how to use parameters to define a file name which will eventually point to a specific file in our dataset.
-
Next up, we have to define the data lake that we want to copy this data to. Navigate to the Sink tab and click new.
-
Select Azure Data Lake Storage Gen 2
- Select Delimited Text
- Input ‘CovidLatest’ in the name field
-
Select new linked service
-
Input fields as shown below:
Name: ADLSLSReference
Azure subscription: MTC Sydney Azure
Storage account name: stcovidhackoutputprod
- Click test connection.
- Next, to set a file path click browse
- Click on data and then select inputs
-
Click OK.
-
Click Validate All
-
If all the validations are correct, click Publish All
-
Next, click Add trigger
-
Click Trigger now to trigger the pipeline straight away. When prompted for a parameter, input a date value in the format of MM-DD-YYYY. Click OK. The pipeline will then start to run.
-
To check how the pipeline is running, go to the monitor step in Azure Data Factory.






























