Open source repositories tagged with #spark, ranked by health score.
Apache Paimon is a lake format that enables building a Realtime Lakehouse Architecture with Flink and Spark for both streaming and batch operations.
Nessie: Transactional Catalog for Data Lakes with Git-like semantics
YTsaurus is a scalable and fault-tolerant open-source big data platform.
Rust based high-performance Apache Uniffle shuffle-server