This is a simple implementation as a POC to showcase the MLflow AI Gateway component in providing a unique endpoint to query/predict different LLM Models.
Two open LLM models like Mistral-ai-7b and Llama-2-7b have been used to run inference and evalute using built-in metrics to compare them.
!!!No API_KEY is Needed
you have:
- python3 > 3.10
- pip is installed
- run
pip install -r requirements.txt
The code uses transformers from huggingface portal
- VM_1: A VM with one GPU RAM: 15GB where you launch the AI gateway deployment
- VM_2: A VM with two GPU RAM: 15GB where you serve the two AI models using "mlflow models serve"
python mistral-ai-7b-quant.py
mlflow server --backend-store-uri postgresql://<db_user_name>:<db_user_password>@0.0.0.0/<mlflow_db_name> --artifacts-destination <path> -h 0.0.0.0
-p 5000 --app-name basic-auth
In the VM_1 ensure that mlflow server is running and the following MLflow config vars are set properly. MLFLOW_TRACKING_URI=https://mlflow.dev.ai4esoc.eu MLFLOW_TRACKING_USERNAME=******** MLFLOW_TRACKING_PASSWORD=********
Ensure mlflow is running:
- llama-LLM:
mlflow models serve -m mlflow-artifacts:/1/5dfaa760ea4740e5a1d44009c3bdffb2/artifacts/artifacts -h 0.0.0.0 -p 8008 -t 1200 --no-conda - mistral-ai-LLM:
mlflow models serve -m mlflow-artifacts:/1/f0df557a9f72498dbf8168da0bffc52b/artifacts/artifacts -h 0.0.0.0 -p 8000 -t 1200 --no-conda
Launch the AI gateway where via a unified URL, the LLM models will be served (first check the llm-config.yaml file)
- set the configuration variable:
export MLFLOW_DEPLOYMENTS_CONFIG=/path/to/config.yaml
export MLFLOW_DEPLOYMENTS_TARGET=http://127.0.0.1:8001
- Launch MLflow server using this command:
mlflow gateway start --config-path /path/to/config.yaml --port 8001
eval_results.json file are generated for each run.
This work is co-funded by AI4EOSC project that has received funding from the European Union's Horizon Europe 2022 research and innovation programme under agreement No 101058593
For open source projects, say how it is licensed.