Skip to content

Repository files navigation

Ariadne

CI Maven Central License Spark Scala

Find the thread through your data lake. Ariadne builds lightweight Delta-backed indexes over your Parquet, CSV, and JSON files so Spark joins read only the files they actually need.

📖 Full documentation →

Install

<!-- Spark 3.5 / Delta 3.2 (Scala 2.12, Java 11) -->
<dependency>
    <groupId>dev.cjfravel</groupId>
    <artifactId>ariadne-spark35_2.12</artifactId>
    <version>0.1.10-beta</version>
</dependency>

<!-- Spark 4.1 / Delta 4.1 (Scala 2.13, Java 21) -->
<dependency>
    <groupId>dev.cjfravel</groupId>
    <artifactId>ariadne-spark41_2.13</artifactId>
    <version>0.1.10-beta</version>
</dependency>

Quick look

import dev.cjfravel.ariadne.Index
import dev.cjfravel.ariadne.Index._  // for df.join(index, ...) implicit

spark.conf.set("spark.ariadne.storagePath",
  s"abfss://$container@$account.dfs.core.windows.net/ariadne")

val index = Index("orders", orderSchema, "parquet")
index.addIndex("customer_id")
index.addFile(orderFiles: _*)
index.update

// Ariadne loads only the files that match
val result = customerDf.join(index, Seq("customer_id"), "inner")

Documentation

Everything lives on the docs site:

Examples

Interactive Jupyter notebooks demonstrating every feature are in examples/. Run them with:

cd examples && docker compose up --build

Then open http://localhost:8888.

Contributing

See the contributor guide and Code of Conduct. All contributions are subject to the Contributor License Agreement.

Citation metadata is available in CITATION.cff.

Security

To report a vulnerability, please follow the security policy — do not open a public issue.

License

Ariadne is licensed under the MIT License with a SaaS provision. See LICENSE.md.

About

A Spark library that builds file-level indexes for efficient data lake joins

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages