Stars
An orchestration platform for the development, production, and observation of data assets.
A dynamic data completeness and accuracy library at enterprise scale for Apache Spark
A PyTorch implementation of Convolutional Sequence Embedding Recommendation Model (Caser)
A Matlab implementation of Convolutional Sequence Embedding Recommendation Model (Caser)
Jenga is an experimentation library that allows data science practititioners and researchers to study the effect of common data corruptions (e.g., missing values, broken character encodings) on the…
Smart Automation Tool for building modern Data Lakes and Data Pipelines
Automated data quality suggestions and analysis with Deequ on AWS Glue
R package implementing the SA-CCR based on the CRR2 Regulation
Deequ is a library built on top of Apache Spark for defining "unit tests for data", which measure data quality in large datasets.