Find the thread through your data lake. Ariadne builds lightweight Delta-backed indexes over your Parquet, CSV, and JSON files so Spark joins read only the files they actually need.
<!-- Spark 3.5 / Delta 3.2 (Scala 2.12, Java 11) -->
<dependency>
<groupId>dev.cjfravel</groupId>
<artifactId>ariadne-spark35_2.12</artifactId>
<version>0.1.10-beta</version>
</dependency>
<!-- Spark 4.1 / Delta 4.1 (Scala 2.13, Java 21) -->
<dependency>
<groupId>dev.cjfravel</groupId>
<artifactId>ariadne-spark41_2.13</artifactId>
<version>0.1.10-beta</version>
</dependency>import dev.cjfravel.ariadne.Index
import dev.cjfravel.ariadne.Index._ // for df.join(index, ...) implicit
spark.conf.set("spark.ariadne.storagePath",
s"abfss://$container@$account.dfs.core.windows.net/ariadne")
val index = Index("orders", orderSchema, "parquet")
index.addIndex("customer_id")
index.addFile(orderFiles: _*)
index.update
// Ariadne loads only the files that match
val result = customerDf.join(index, Seq("customer_id"), "inner")Everything lives on the docs site:
- Getting Started — install, configure, build your first index
- Index Types — regular, bloom, range, temporal, computed, exploded
- Usage Guide — joins, column selection, Spark SQL catalog
- Configuration — every
spark.ariadne.*setting - Maintenance — compact, vacuum, delete files
- Troubleshooting — exceptions and common fixes
- Architecture — internals for contributors
Interactive Jupyter notebooks demonstrating every feature are in examples/. Run them with:
cd examples && docker compose up --buildThen open http://localhost:8888.
See the contributor guide and Code of Conduct. All contributions are subject to the Contributor License Agreement.
Citation metadata is available in CITATION.cff.
To report a vulnerability, please follow the security policy — do not open a public issue.
Ariadne is licensed under the MIT License with a SaaS provision. See LICENSE.md.
