Polish-Cave-Data-Scraper is a robust Python-based tool designed to scrape and collect comprehensive data on Polish caves from the Central Geological Database of Polish Caves (CBDG) managed by the Polish Society for Friends of Earth Sciences (PTPNoZ). The scraper gathers standardized information, including geolocation, morphology, environmental data, historical descriptions, and graphic attachments such as plans, sections, and photographs. This dataset serves as a valuable resource for researchers, conservationists, and speleologists interested in the geological and environmental aspects of Polish caves.
- Python 3.9 or higher
- uv (Python package and environment manager)
-
First, ensure you have uv installed on your system. For example:
curl -LsSf https://astral.sh/uv/install.sh | sh -
Clone the repository:
git clone https://github.com/yourusername/polish-cave-data-scraper.git cd polish-cave-data-scraper -
Install project dependencies using uv:
uv sync
To ensure a clean environment for the project:
- Reinstall the locked environment:
uv sync --reinstall
The scraper consists of three main scripts that should be run in sequence:
-
First, run the data fetching script:
uv run python fetch.py
This script collects raw data from the CBDG database.
-
Then, run the parsing script:
uv run python parse.py
This script processes the collected data into a structured format.
-
Finally, run the cleaning script:
uv run python clean.py
This script transforms and cleans the data using PySpark.
The project includes a separate script for downloading bibliography data from the CBDG database:
uv run python download_bibliography.pyThis script fetches all bibliography records from the Polish Geological Institute's cave database using the JSON endpoint. The bibliography includes citations for publications related to Polish caves, organized by region.
Features:
- Downloads complete bibliography dataset via paginated API
- Filters by author, year, title, or cave region
- Saves data to
bibliography.jsonlin JSON Lines format (matching project conventions) - Automatically trims whitespace from all string fields
- Handles session cookies and request headers automatically
The script can be customized by editing configuration variables in the main() function for specific search criteria.
Cave images (plans, sections, and diagrams) can be upscaled and denoised using waifu2x-ncnn-vulkan:
-
Download waifu2x-ncnn-vulkan from releases and extract to project directory
-
Run the upscaling script:
uv run python upscale_images.py
This processes all images in caves/ directory, applying 2x upscaling and level-2 denoising. Upscaled images are saved to caves_upscaled/ with the same directory structure.
The project uses modern Python code quality tools to maintain high standards:
- ruff - Fast linting, import sorting, code modernization, and formatting
- ty - Fast static type checking
- pytest - Test runner
- pre-commit - Git hooks for automated quality checks
# Run linter and auto-fix issues
uv run ruff check . --fix
# Format code
uv run ruff format .
# Check types
uv run ty check
# Run tests
uv run pytest
# Run all checks at once
uv run pre-commit run --all-filesInstall pre-commit hooks to automatically run quality checks before each commit:
# One-time setup
uv run pre-commit install
# Run manually on all files
uv run pre-commit run --all-filesThe hooks are local and run through uv run, so they use the same locked project environment as normal development commands.
uv run pytest