Skip to content

About

Federated learning server app for fine-tuning LLMs with distributed clients. Build using Flower / FlowerTune. The FL server is deployed from the AI4EOSC dashboard.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

EOSC ARENA Flower LLM Server

This repository contains a Flower-based Federated LLM server deployment intended to run on EOSC ARENA. Clients (SuperNodes) run on separate nodes and connect to the SuperLink (server) deployed on the platform.

Overview

  • Server (SuperLink): runs the Flower SuperLink process that accepts connections from SuperNodes. The provided script to start it is start_superlink.sh.
  • Clients (SuperNodes): run on separate machines, each with its own local dataset. Clients connect to the deployed server and perform local training.

Step-by step workflow:

  1. Ensure the Flower CLI config file exists and points the exec API address to the SuperLink exec API. Verify the config path with:
flwr config list

This prints the active Flower config file path (look for the line labeled Flower Config file).

  1. The file deploy/flwr_config.toml is used for Flower CLI connections. If the exec API address or TLS settings need to change, edit deploy/flwr_config.toml, not ~/.flwr/config.toml directly.

  2. Copy deploy/flwr_config.toml to $HOME/.flwr/config.toml, by default /root/.flwr/config.toml in the server container:

chmod +x ./install_flwr_config.sh
./install_flwr_config.sh

Note that must be done only in the first run or if you need to re-start the app.

  1. Start the SuperLink on the server side (this script expects Traefik routing to terminate TLS):
cd arena-fl-server-llm
./start_superlink.sh

If you changed the config file after starting the SuperLink, stop and restart the SuperLink so any new flwr run use the updated configuration.

Client-side (SuperNode) workflow:

  1. On each client machine create and activate a Python virtual environment (we recommend using Python 3.12):
python3.12 -m venv .venv
source .venv/bin/activate
  1. Install dependencies using the same pyptroject.toml file as in the server side:
pip install -e .
  1. Each client must have its own local dataset (never centrally shared). Then run the SuperNode start script with the server route, the path to the local dataset, and the local port the SuperNode will expose. Example:
./start_supernode.sh fedserver-<DEPLOYMENT_UUID>.<DATA_CENTER>-deployments.cloud.ai4eosc.eu <DATA_PATH> <PORT>
  • The last argument is the local port used by the SuperNode and can be changed if needed (make sure it does not conflict with other services on the same machine).
  • Client code/structure example: ai4os/arena-fl-client-llm.

Starting a remote run and checking status

After the SuperLink is running on the server side and the clients (SuperNodes) have connected, start the federated run from the server container using the Flower CLI:

cd arena-fl-server-llm
flwr run . remote

Then, you can check the runs:

flwr ls remote

If you want to check the logs of an specific run:

flwr log <RUN_ID>

Environment variables that can be customized in the dashboard:

Variable Default
NUM_ROUNDS 10
MODEL_NAME Qwen/Qwen2.5-0.5B
MODEL_QUANTIZATION 4
NUM_EPOCHS 3
FRACTION_TRAIN 0.2
FRACTION_EVALUATE 0.0

Warning

This project is under active development.

License

This project is licensed under the Apache 2.0 license.

Funding and acknowledgments

This work is funded by European Union through the EOSC-ARENA project (Horizon Europe) under Grant number 101292597.

About

Federated learning server app for fine-tuning LLMs with distributed clients. Build using Flower / FlowerTune. The FL server is deployed from the AI4EOSC dashboard.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages