This is the repository associated to the review by Luna Liviero, Philipp Leclercq, Erwan Privat, Sergio Rampino and Antonino Polimeno under review in the Sustainability and Circularity NOW journal.
The motivation and context for this work is described in the aforementioned article. We focus here only on the illustrative example results of accuracy improvement by combining the datasets from Odegova et al.1 and Luu et al.2 Available code here is based on the first1 team code. Repository links are in the bibliography.
More information on the datasets can be found on the main review text. To summarize, the 2024 publication by Odegova et al.1 compiled a dataset of DES properties from experimental studies dating back to 2003, resulting in 2,303 data entries for melting temperature, 4,369 for density, and 4,216 for viscosity. In 2023, Luu et al.2 collected a dataset with 402 entries from literature, reporting the melting temperature as well as the component percentages for each solvent, and successfully developed a prompt-based transformer model for both predictive and generative tasks, using SMILES to represent the molecules. They trained the model jointly on DES tasks and tasks related to the well-established and much larger QM93 dataset.
data/DES_Melting.csv: Odegova et al.1 dataset,data/DES1.csv: Luu et al.2 dataset,data/DES_TMELT.csv: our combined dataset,data/DES_TMELT.ipynb: code used for the merging.
For illustrative purposes, and building on the data available from
recent works on DESs, computational analysis was conducted leading to
insightful exploration in regard to two aspects: the impact of
dataset size and the contribution of feature selection to model
prediction performance. The datasets compiled by Luu et al.2 and
Odegova et al.1, along with the latter’s codebase, served as
the foundation for further investigation. The first step involved
compiling a unified dataset by merging the existing data frames. Luu
et al.2 provides a file with 401 entries, while Odegova at
al.1 contributed 2,259 entries. The data is homogenized
through unit conversions where necessary and by retaining only the
common columns. Eight redundant entries were identified and handled
differentially. For seven of these, both data pairs were retained, as
they showed slight variations in component melting
temperatures—considered acceptable within a ±10 K margin. These
duplicates were grouped by averaging the melting temperature values
to preserve as much information as possible with some curating. After
merging, the resulting dataset contained 2,651 entries. Each DES
entry includes, where available: the SMILES formula of both
components, the melting temperature of each component, molar ratios,
the melting temperature of the resulting DES, the full names of both
components, the DES type, and the reference/DOI of the source study.
From this dataset, the final feature set and label were extracted for
model training and testing. Using the codebase from Odegova et
al.1, molecular descriptors were generated from SMILES
formulas using the RDKit library.4 Feature selection followed
the methodology in the reference article, where over 200 descriptors
were filtered using correlation matrices to exclude highly correlated
variables. The final features included singular melting temperature,
molar fractions, molecular weight, hydrogen bond donor count, various
functional group counts, and toxicity. The target label was the
melting temperature of the DES. The models are trained on 80 % of
entries and tested on 20 % of them, using 5-fold cross-validation.
The “mixtures-out”5 method was applied to prevent data
leakage by grouping entries containing the same molecule into the
same subset. Hyperparameter optimization was performed for each
model. Evaluation metrics included the coefficient of determination
(
The results indicate that the relevance of the "type of DES" feature
varies across models. The kernel based models, SVM, KNN and MLP, were
the most sensitive to this feature removal with a notable drop in
performance. In contrast, having more samples always leads to
improved performance across all models. Also in this case, the kernel
models are deeply affected in their performance. This highlights the
importance of quality, structured, abundant data in the field of
machine learning to enhance model accuracy. One limitation of this
analysis lies in the uneven distribution of DES classification.
Although the additional entries lacked class information, it is
assumed that their inclusion did not attenuate the imbalance, leaving
type III and V disproportionately represented. It cannot be excluded
that the seemingly improved model performance is driven primarily by
the overrepresented classes, while the model may struggle to capture
relationships within underrepresented DES types. Despite these
concerns, the model's performance remains significant, with three
models exceeding an
In the table below is a recap of the cross-validation metrics
| Model | Original |
Extended |
Original RMSE | Extended RMSE |
|---|---|---|---|---|
| DTR | 0.583633 | 0.652694 | 47.3243 | 47.7102 |
| RFR | 0.713469 | 0.785704 | 39.3110 | 37.2780 |
| GBR | 0.733936 | 0.788667 | 37.9431 | 36.6524 |
| CBR | 0.757018 | 0.823381 | 36.2145 | 33.8333 |
| XGB | 0.733484 | 0.803071 | 38.0284 | 35.6826 |
| SVM | 0.629013 | 0.812928 | 43.7242 | 34.8591 |
| KNN | 0.642345 | 0.805520 | 43.0502 | 35.3635 |
| MLP | 0.469012 | 0.771632 | 52.9626 | 38.4747 |
Regression models:
- DTR: decision tree
- RFR: random forest
- GBR: gradient boosting
- CBR: cat boosting
- XGB: extreme boosting
- SVM: support vector machine
- KNN: K-nearest neighbors
- MLP: multilayer perceptron
The source can be found in the src directory. Please note that
most of the code is an adaptation of the already existing one made
by Odegova at al.1, link to repository in bibliography.
Footnotes
-
Odegova, V.; Lavrinenko, A.; Rakhmanov, T.; Sysuev, G.; Dmitrenko, A.; Vinogradov, V. Green Chem., 2024, 26, 3958. https://github.com/lamm-mit/MoleculeDiffusionTransformer ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
-
Luu, R.K.; Wysokowski, M.; Buehler, M.J. Appl. Phys. Lett. 2023, 122, 234103. https://github.com/lamm-mit/MoleculeDiffusionTransformer ↩ ↩2 ↩3 ↩4 ↩5
-
Junde Li, Swaroop Ghosh (2024). Dataset: QM9. https://doi.org/10.57702/exp0m45r ↩
-
RDKit: Open-source cheminformatics. https://www.rdkit.org ↩
-
Muratov, E. N., Varlamova, E. V., Artemenko, A. G., Polishchuk, P. G., & Kuz'min, V. E. (2012). Existing and developing approaches for QSAR analysis of mixtures. Molecular informatics, 31(3‐4), 202-221. ↩
-
Petteri Vainikka, Sebastian Thallmair, Paulo Cesar Telles Souza, and Siewert J. Marrink, ACS Sustainable Chemistry & Engineering 2021 9 (51), 17338-17350, DOI: 10.1021/acssuschemeng.1c06521 ↩