Skip to content

Repository files navigation

Getting started

This is a simple implementation as a POC to showcase the MLflow AI Gateway component in providing a unique endpoint to query/predict different LLM Models. Two open LLM models like Mistral-ai-7b and Llama-2-7b have been used to run inference and evalute using built-in metrics to compare them.
!!!No API_KEY is Needed

Prerequisites

Main installation

you have:

  • python3 > 3.10
  • pip is installed

Install dependencies from the requirements.txt

  • run pip install -r requirements.txt

The code uses transformers from huggingface portal

Hardware requirements to execute the .py files

  • VM_1: A VM with one GPU RAM: 15GB where you launch the AI gateway deployment
  • VM_2: A VM with two GPU RAM: 15GB where you serve the two AI models using "mlflow models serve"

Execute one-by-one the .py files, e.g. mistral-ai-7b-quant.py

python mistral-ai-7b-quant.py

Launch MLflow tracking server in VM_1

mlflow server --backend-store-uri postgresql://<db_user_name>:<db_user_password>@0.0.0.0/<mlflow_db_name> --artifacts-destination <path> -h 0.0.0.0  
-p 5000 --app-name basic-auth

In the VM_1 ensure that mlflow server is running and the following MLflow config vars are set properly. MLFLOW_TRACKING_URI=https://mlflow.dev.ai4esoc.eu MLFLOW_TRACKING_USERNAME=******** MLFLOW_TRACKING_PASSWORD=********

Deploy the two LLM models using "mlflow models serve" command (in VM_2)

Ensure mlflow is running:

  • llama-LLM: mlflow models serve -m mlflow-artifacts:/1/5dfaa760ea4740e5a1d44009c3bdffb2/artifacts/artifacts -h 0.0.0.0 -p 8008 -t 1200 --no-conda
  • mistral-ai-LLM: mlflow models serve -m mlflow-artifacts:/1/f0df557a9f72498dbf8168da0bffc52b/artifacts/artifacts -h 0.0.0.0 -p 8000 -t 1200 --no-conda

Launch the AI gateway where via a unified URL, the LLM models will be served (first check the llm-config.yaml file)

  • set the configuration variable:
export MLFLOW_DEPLOYMENTS_CONFIG=/path/to/config.yaml 
export MLFLOW_DEPLOYMENTS_TARGET=http://127.0.0.1:8001
  • Launch MLflow server using this command:
mlflow gateway start --config-path /path/to/config.yaml --port 8001

Check the evaluation results after each executions

eval_results.json file are generated for each run.

Acknowledgment

This work is co-funded by AI4EOSC project that has received funding from the European Union's Horizon Europe 2022 research and innovation programme under agreement No 101058593

License

For open source projects, say how it is licensed.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages