-
Create a Python 3.9 virtual environment. Several of the pinned dependencies (
scikit-learn 0.24.2,scipy 1.7.1,matplotlib 3.4.3,h5py 3.2.1) only publish cp38/cp39 wheels, so a newer interpreter forces a source build that does not succeed. -
Download/clone dynamics-toolbox and install it first, into that same virtual environment:
git clone https://github.com/LucasCJYSDL/dynamics-toolbox.git cd dynamics-toolbox pip install -r requirements.txt pip install -e . cd ..
Order matters.
dynamics-toolboxpins its own versions oftorch,pytorch-lightning,torchmetrics,hydra-core,scikit-learn,matplotlib,h5pyandscipy; whichever of the two requirement sets is installed second silently overwrites the other. This folder'srequirements.txtis aligned with those pins, so installing it second lets pip surface any remaining incompatibility as an error instead of a downgrade. -
Then install this folder's dependencies and build the Cython extension:
pip install -r requirements.txt cd offlinerlkit/utils/ctree && python setup.py build_ext --inplace && cd ../../..
The
ctreebuild is required for every algorithm, not only the search-based ones:offlinerlkit/policy/__init__.pyimportsbamcts.py, which imports the compiledsearch_treeextension.
Before running anything, edit the paths at the top of
rl_preparation/process_raw_data.py. They are hard-coded to the CMU cluster
layout and will not exist anywhere else:
raw_data_dir = "/zfsauton/project/fusion/data/organized/noshape_gas_flattop_synthesized"
training_model_dir = "/zfsauton/project/fusion/models/rpnn_noshape_gas_flat_top_step_two_logvar"
evaluation_model_dir = "/zfsauton/project/fusion/models/rpnn_noshape_gas_flat_top_step_two_logvar"dynamics/synthesize_rollouts.py contains a second hard-coded raw_data_dir
that needs the same treatment. Note that these absolute paths are read directly —
the code does not look for a Fusion/data/ directory.
-
Please start from converting the raw fusion data to the format required by offline RL or Imitation Learning:
python rl_preparation/process_raw_data.py
-
You can run different offline RL algorithms simply by:
python rl_scripts/run_XXX.py --task YYY --seed Z python rl_scripts/run_rombrl.py --task betan --seed 0
- XXX can be one of [rombrl, cql, edac, combo, mobile, bamcts, rambo], corresponding to our algorithm and the 6 baselines used in the paper.
- YYY can be one of [betan, dens_component1, rotation_component1], corresponding to the tracking tasks for betan, density, and rotation, respectively. The value is matched as a name prefix against the state list in
rl_preparation/state_actuator_spaces.py. - The random seed Z can be one of [0, 1, 2].
--envselects the environment wrapper, one of [base, profile_control]; it defaults toprofile_control.
Unlike the scripts under
D4RL/, these read--taskand--seedfrom the command line as expected.
- Plots of episodic returns during the policy learning process, along with trained policy model checkpoints, will be generated and saved in the 'log' folder.
The execution of these experiments relies on operational data from DIII-D, which is protected. As a result, we are unable to release them until we obtain the necessary approvals.