diff --git a/README.md b/README.md index 115a0ed..17b4923 100644 --- a/README.md +++ b/README.md @@ -16,8 +16,7 @@ use the FairNow API's. The recommended order is: 1. `Getting Started` - demonstrates how to configure authorization and make a first API call. 2. `Applications API` - create and read an `AI Application` with the API. `AI Applications` are a core building block for the FairNow system. - 3. `Generating Synthetic User Bias Test Data` - construct a synthetic test data file to be used to simulate ML model bias evaluation. Constructing synthetic data is optonal. - 4. `User Data Testing` - how to perform ML model bias testing via the API. + 3. `Export Reports` - generate tsvs for application and company compliance, inventory, and risk and severity ### Note to contributors: diff --git a/notebooks/Applications API.ipynb b/notebooks/Applications API.ipynb index 9725540..2477bc1 100644 --- a/notebooks/Applications API.ipynb +++ b/notebooks/Applications API.ipynb @@ -114,7 +114,7 @@ "id": "8", "metadata": {}, "source": [ - "#### Save the `application_id` so that you can perform other compliance tasks for the application like approving it for deployment, testing for biases or adding evidence that a control has been met." + "#### Save the `application_id` so that you can perform other compliance tasks for the application like approving it for deployment or adding evidence that a control has been met." ] }, { diff --git a/notebooks/Generating Synthetic User Bias Test Data.ipynb b/notebooks/Generating Synthetic User Bias Test Data.ipynb deleted file mode 100644 index d210538..0000000 --- a/notebooks/Generating Synthetic User Bias Test Data.ipynb +++ /dev/null @@ -1,141 +0,0 @@ -{ - "cells": [ - { - "cell_type": "markdown", - "id": "0", - "metadata": {}, - "source": [ - "# Generating Synthetic User Bias Test Data\n", - "\n", - "\n", - "#### This notebook demonstrates how to generate synthetic test data to simulate the evaluation of bias for protected classes. This is optional as you may have real scoring data to be evaluated.\n", - "\n", - "### Data Requirements\n", - "\n", - "#### The CSV data files require the first row to contain the column names.\n", - "\n", - "#### The following columns are required:\n", - "* `TimeStamp` (ISO8601 Timestamp, e.g `2023-12-14T16:26:05.898156Z`)\n", - "* `Score` (a number between 0 and 1)\n", - "* `Protected Class Columns` (race, gender, disability_status, gender_identity)\n", - "\n", - "#### Optionally, you can add up to 3 additional columns to use for grouping and filtering the bias results after testing." - ] - }, - { - "cell_type": "markdown", - "id": "5", - "metadata": {}, - "source": [ - "#### Now generate some random data. In addition to the required columns `Timestamp` and `Score` and the protected class columns the example contains two columns that can be used to filter the results. " - ] - }, - { - "cell_type": "code", - "execution_count": 7, - "id": "6", - "metadata": {}, - "outputs": [], - "source": [ - "import datetime\n", - "import random\n", - "import csv\n", - "\n", - "# The default protected class columns will be used if none are provided.\n", - "# Update protected class columns by updating the list below:\n", - "protected_class_columns = [\"race\", \"gender\", \"disability_status\", \"gender_identity\"]\n", - "\n", - "number_of_protected_class_values = 10\n", - "number_of_rows = 1000\n", - "\n", - "required_columns = [ \"Timestamp\", \"Score\" ]\n", - "filter_columns = [ \"Color\", \"Size\"]\n", - "all_columns = required_columns + protected_class_columns + filter_columns\n", - "\n", - "colors = [ \"Blue\", \"Red\", \"Yellow\", \"Green\", \"Orange\", \"Purple\" ]\n", - "sizes = [ \"X-Small\", \"Small\", \"Medium\", \"Large\", \"X-Large\" ]\n", - "\n", - "data = [ all_columns ]\n", - "for x in range(number_of_rows):\n", - " row = [\n", - " datetime.datetime.now(datetime.timezone.utc).isoformat().replace(\"+00:00\", \"Z\"), # ISO Timestamp\n", - " random.random() # Score\n", - " ]\n", - " for column in protected_class_columns:\n", - " row.append(f\"{column}-{random.randrange(number_of_protected_class_values)}\")\n", - " \n", - " row.append(random.choice(colors)) \n", - " row.append(random.choice(sizes))\n", - " \n", - " data.append(row)\n" - ] - }, - { - "cell_type": "markdown", - "id": "7", - "metadata": {}, - "source": [ - "#### Now create the CSV file in the local directory." - ] - }, - { - "cell_type": "code", - "execution_count": 8, - "id": "8", - "metadata": {}, - "outputs": [], - "source": [ - "with open('scores.csv', 'w', newline='') as csvfile:\n", - " writer = csv.writer(csvfile)\n", - " writer.writerows(data)\n" - ] - }, - { - "cell_type": "markdown", - "id": "9", - "metadata": {}, - "source": [ - "#### Check that the file exists and contains the randomly generated data.\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "10", - "metadata": {}, - "outputs": [], - "source": [ - "!cat scores.csv" - ] - }, - { - "cell_type": "markdown", - "id": "11", - "metadata": {}, - "source": [ - "#### You can now use this file for testing bias - see the `User Data Testing` notebook for instructions." - ] - } - ], - "metadata": { - "kernelspec": { - "display_name": ".venv", - "language": "python", - "name": "python3" - }, - "language_info": { - "codemirror_mode": { - "name": "ipython", - "version": 3 - }, - "file_extension": ".py", - "mimetype": "text/x-python", - "name": "python", - "nbconvert_exporter": "python", - "pygments_lexer": "ipython3", - "version": "3.13.7" - } - }, - "nbformat": 4, - "nbformat_minor": 5 -} diff --git a/notebooks/User Data Testing.ipynb b/notebooks/User Data Testing.ipynb deleted file mode 100644 index 6b98394..0000000 --- a/notebooks/User Data Testing.ipynb +++ /dev/null @@ -1,289 +0,0 @@ -{ - "cells": [ - { - "cell_type": "markdown", - "id": "0", - "metadata": {}, - "source": [ - "## Running FairNow's User Data Bias Testing" - ] - }, - { - "cell_type": "markdown", - "id": "1", - "metadata": {}, - "source": [ - "#### FairNow's User Data Testing is a way to evaluate a model for bias using real data. Please do not send any PII data." - ] - }, - { - "cell_type": "markdown", - "id": "2", - "metadata": {}, - "source": [ - "### Prerequisites:\n", - "\n", - "#### To use this notebook, you'll need a `Client ID` and `Client Secret`. These will either have been provided to you, or you can generate from https://app.fairnow.ai and going the the Admin menu. This notebook assumes you have these available to enter when prompted.\n", - "\n", - "#### To run the simulation you will need an `application_id` for the specific model you want to test. Details of how to create and lookup AI Applications can be found in the `Applications API` notebook.\n", - "\n", - "#### Finally, you'll need a `threshold` value, which is the value at which anyone with a score above is considered a passing score.\n", - "\n", - "#### Running the following cell will prompt you for the `Client ID` and `Client Secret` and create a `client` instance that can be used to communicate with the Fairnow APIs.\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "3", - "metadata": {}, - "outputs": [], - "source": [ - "import json\n", - "\n", - "from utils.fairnow import get_client\n", - "\n", - "client = get_client(client_id='client_id') # Replace with your Client Id" - ] - }, - { - "cell_type": "markdown", - "id": "4", - "metadata": {}, - "source": [ - "#### You will also need to prepare a CSV file to upload containing data. The first row contains the column names.\n", - "\n", - "#### The following columns are required:\n", - "* `TimeStamp` (ISO8601 Timestamp, e.g `2023-12-14T16:26:05.898156Z`)\n", - "* `Score` (a number between 0 and 1)\n", - "* `Each of the Protected Class Columns` (see `Generating User Bias Test Data` notebook for example on how to lookup the column names and example for generating a CSV to upload.)\n", - "\n", - "#### Additional columns can be added to allow filtering of data, e.g. `Job Title`, `Location` etc\n", - "\n", - "### Place the CSV file in the same directory as this notebook." - ] - }, - { - "cell_type": "code", - "execution_count": 3, - "id": "5", - "metadata": {}, - "outputs": [], - "source": [ - "scores_file_name = 'scores.csv' # Change the filename if different" - ] - }, - { - "cell_type": "markdown", - "id": "6", - "metadata": {}, - "source": [ - "#### Next we create a test. The test needs an `application_id`, along with a threshold value. The response will include the `test_id` and a pre-signed URL used upload the CSV file." - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "7", - "metadata": {}, - "outputs": [], - "source": [ - "application_id = '' # Replace with your own Application Id\n", - "threshold = 0.5 # Change to your own threshold setting\n", - "protected_class_columns = ['race', 'gender', 'disability_status', 'gender_identity'] # Default protected class columns\n", - "\n", - "if not application_id:\n", - " raise ValueError(\"Please set application_id\")\n", - "\n", - "if threshold <= 0.0 or threshold >= 1.0:\n", - " raise ValueError(\"Threshold must be between 0.0 and 1.0\")\n", - "\n", - "if not protected_class_columns:\n", - " raise ValueError(\"Please set protected_class_columns. Default is ['race', 'gender', 'disability_status', 'gender_identity']\")\n", - "\n", - "start_test_route = f\"/applications/{application_id}/tests\"\n", - "\n", - "\n", - "test_type = \"fairness_ml_user\" # Do not change test_type\n", - "\n", - "request_body = {\n", - " \"test_name\": \"API Client Testing\",\n", - " \"test_description\": \"Testing User Bias\",\n", - " \"test_type\": test_type,\n", - " \"inputs\": {\n", - " \"type\": test_type,\n", - " \"threshold\": threshold,\n", - " \"protected_class_columns\": protected_class_columns\n", - " }\n", - "}\n", - "\n", - "response = client.post(start_test_route, json=request_body, timeout=None)\n", - "\n", - "if response.status_code == 200:\n", - " print(\"New Test has been created:\")\n", - " print(json.dumps(response.json(), indent=4))\n", - "else:\n", - " print(f\"API Error Response: {response.status_code} - {response.text}\")\n" - ] - }, - { - "cell_type": "markdown", - "id": "8", - "metadata": {}, - "source": [ - "#### Use the pre-signed upload url to upload the scores file. Once the scores are uploaded, this triggers the analysis job. This runs in the background again and can take a few minutes." - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "9", - "metadata": {}, - "outputs": [], - "source": [ - "import httpx\n", - "\n", - "response_body = response.json()\n", - "test_id = response_body[\"test_id\"]\n", - "presigned_upload = response_body[\"presigned_url_scores_upload\"]\n", - "upload_url = presigned_upload[\"url\"]\n", - "upload_fields = presigned_upload[\"fields\"]\n", - "upload_key = upload_fields[\"key\"]\n", - "\n", - "data_file = {'file': (upload_key, open(scores_file_name, 'rb'))}\n", - "\n", - "# Upload the scores CSV file\n", - "response = httpx.post(upload_url, data=upload_fields, files=data_file)\n", - "\n", - "if response.status_code == 204:\n", - " print(\"File has been uploaded.\")\n", - "else:\n", - " print(f\"Error uploading file: {response.status_code} - {response.text}\")" - ] - }, - { - "cell_type": "markdown", - "id": "10", - "metadata": {}, - "source": [ - "#### We'll query the API again to know when the analysis has been finished" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "11", - "metadata": {}, - "outputs": [], - "source": [ - "import time\n", - "\n", - "test_route = f\"/applications/{application_id}/tests/{test_id}\"\n", - "\n", - "response = client.get(test_route)\n", - "if response.status_code == 200:\n", - " current_status = response.json()[\"status\"][\"id\"]\n", - " print(f\"Current test status: {current_status}\")\n", - "else:\n", - " raise ValueError(f\"API Error Response: {response.status_code} - {response.text}\")\n", - "\n", - "while current_status not in ['ready', 'error']:\n", - " time.sleep(15)\n", - " response = client.get(test_route)\n", - " if response.status_code == 200:\n", - " current_status = response.json()[\"status\"][\"id\"]\n", - " print(f\"Current test status: {current_status}\")\n", - " else:\n", - " raise ValueError(f\"API Error Response: {response.status_code} - {response.text}\")\n", - "\n", - "if current_status == 'error':\n", - " error_details = response.json()[\"file_validation_report\"]\n", - " print('Analysis encountered an error. Details:')\n", - " print(error_details)\n", - "else:\n", - " print(f'Analysis results ready to download.')\n" - ] - }, - { - "cell_type": "markdown", - "id": "12", - "metadata": {}, - "source": [ - "#### Now the the testing is complete you can check the result in the Fairnow application." - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "13", - "metadata": {}, - "outputs": [], - "source": [ - "test_results_url = response.json()[\"results_url\"]\n", - "print(f\"Check your results in the Fairnow application at '{test_results_url}\")" - ] - }, - { - "cell_type": "markdown", - "id": "14", - "metadata": {}, - "source": [ - "#### Optionally, you can download the raw analysis results with presigned link. The output is a csv file with the results of bias by protected classes." - ] - }, - { - "cell_type": "code", - "execution_count": 17, - "id": "15", - "metadata": {}, - "outputs": [], - "source": [ - "presigned_test_results_url = response.json()[\"presigned_url_test_results\"]\n", - "\n", - "response = httpx.get(presigned_test_results_url)\n", - "\n", - "with open('results.csv', 'wb') as file:\n", - " file.write(response.content)" - ] - }, - { - "cell_type": "markdown", - "id": "16", - "metadata": {}, - "source": [ - "#### Now read the results from the file." - ] - }, - { - "cell_type": "code", - "execution_count": null, - "id": "17", - "metadata": {}, - "outputs": [], - "source": [ - "!cat results.csv" - ] - } - ], - "metadata": { - "kernelspec": { - "display_name": ".venv", - "language": "python", - "name": "python3" - }, - "language_info": { - "codemirror_mode": { - "name": "ipython", - "version": 3 - }, - "file_extension": ".py", - "mimetype": "text/x-python", - "name": "python", - "nbconvert_exporter": "python", - "pygments_lexer": "ipython3", - "version": "3.13.7" - } - }, - "nbformat": 4, - "nbformat_minor": 5 -}