Skip to content

About

Python library to create customized subsets of the MultiCaRe clinical case dataset

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

🌐 multiversity

multiversity is a Python library for creating customized subsets of the MultiCaRe Dataset — an open-source multimodal clinical case dataset with 98,000+ clinical cases and 139,000+ medical images.

PyPI version License: CC0


🚀 Installation

pip install multiversity

🗂️ What is MultiCaRe?

The MultiCaRe Dataset is an open-source collection of de-identified clinical cases and medical images built from PubMed Central open-access case reports. It includes:

  • 98,000+ clinical cases across multiple medical specialties
  • 139,000+ labeled medical images
  • A hierarchical image taxonomy with 140+ classes
  • Rich metadata: patient demographics, clinical narratives, image captions, and citations

✅ Quick Start

1. Import and initialize

from multiversity.multicare_dataset import MedicalDatasetCreator

# Downloads the dataset from Zenodo on first run (5–10 min)
mdc = MedicalDatasetCreator(directory='medical_datasets')

2. Define filters

filters = [
    {'field': 'min_age', 'string_list': ['18']},
    {'field': 'gender', 'string_list': ['Male']},
    {'field': 'case_strings', 'string_list': ['tumor', 'cancer'], 'operator': 'any'},
    {'field': 'label', 'string_list': ['mri', 'head']}
]

3. Create your dataset

mdc.create_dataset(
    dataset_name='brain_tumor_dataset',
    filter_list=filters,
    dataset_type='multimodal'  # Options: 'multimodal', 'text', 'image', 'case_series'
)

4. Explore an example

mdc.display_example()

🔧 Filter Fields

Field Description Example
min_age Minimum patient age ['18']
max_age Maximum patient age ['65']
gender Patient gender ['Male'], ['Female']
case_strings Keywords in the clinical narrative ['tumor', 'cancer']
caption Keywords in image captions ['mass', 'lesion']
label Image taxonomy labels ['mri', 'head']

Use 'operator': 'any' to match any keyword, or 'operator': 'all' (default) to require all.


📦 Dataset Types

Type Description
multimodal Cases with both text and images
text Cases with clinical narratives only
image Image classification dataset
case_series Groups of related cases

📚 Related Resources


🤓 How to Cite

If you use this library or the MultiCaRe dataset, please cite:

Nievas Offidani, M., Roffet, F., González Galtier, M. C., Massiris, M., & Delrieux, C. (2025).
An Open-Source Clinical Case Dataset for Medical Image Classification and Multimodal AI Applications.
Data, 10(8), 123. https://doi.org/10.3390/data10080123

🤝 Contributing

Contributions are welcome! Open an issue or submit a pull request.

If you find this useful, consider giving it a ⭐ — it helps with visibility.

For questions or collaborations, reach out on LinkedIn.

About

Python library to create customized subsets of the MultiCaRe clinical case dataset

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages