Stars
Source code for the X Recommendation Algorithm
Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark
ZIO — A type-safe, composable library for async and concurrent programming in Scala
Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.
Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.
TransmogrifAI (pronounced trăns-mŏgˈrə-fī) is an AutoML library for building modular, reusable, strongly typed machine learning workflows on Apache Spark with minimal hand-tuning
Apache Spark to Apache Cassandra connector
Byzer (former MLSQL): A low-code open-source programming language for data pipeline, analytics and AI.
TiSpark is built for running Apache Spark on top of TiDB/TiKV
This repository contains the development code for sparkMeasure, an Apache Spark performance analysis and troubleshooting library. It simplifies collecting, aggregating, and exporting Spark task/sta…
Essential Spark extensions and helper methods ✨😲
Data Lineage Tracking And Visualization Solution
A simplified, lightweight ETL Framework based on Apache Spark
Apache Spark testing helpers (dependency free & works with Scalatest, uTest, and MUnit)
Spark ClickHouse Connector build on DataSourceV2 API
A Spark SQL extension which provides SQL Standard Authorization for Apache Spark | This repo is contributed to Apache Kyuubi | 项目已迁移至 Apache Kyuubi
Extended datasource support for Spark/Hadoop on Aliyun E-MapReduce.
SparkCube is an open-source project for extremely fast OLAP data analysis. SparkCube is an extension of Apache Spark.
A simplified, lightweight ETL pipeline framework for build stream/batch processing applications on top of Apache Spark
智能数据探索服务(Intelligent Data Exploration Service),一站式Data + AI数据解决方案!
On the fly, translation of Spark programs to run natively on your Oracle DB. Your Spark programs require no changes.