Lists (1)
Sort Name ascending (A-Z)
Stars
This is an entire environment created to master the craft of Spark Declarative Pipelines.
MCP Server for Korean stock analysis. 한국 주식 분석을 위한 MCP 서버입니다.
The Metadata Driven framework for Databricks Lakeflow Declarative Pipelines (formerly Delta Live Tables). Metadata framework that generates production ready Pyspark code for Lakeflow Declarative Pi…
This repo provides a customizable stack for starting new ML projects on Databricks that follow production best-practices out of the box.
Jumpstart CICD deployments in Microsoft Fabric
Extremely fast Query Engine for DataFrames, written in Rust
Apache Arrow is the universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics
Knowledge-based, Content-based and Collaborative Recommender systems are built on MovieLens dataset with 100,000 movie ratings. These Recommender systems were built using Pandas operations and by f…
This repository has moved into https://github.com/dbt-labs/dbt-adapters
Databricks Implementation of the TPC-DI Specification using Traditional Notebooks and/or Delta Live Tables
Metadata driven Spark Declarative Pipelines framework for bronze/silver pipelines
Streaming Synthetic Sales Data Generator: Streaming sales data generator for Apache Kafka, written in Python
A python SPark ETL libRary (SPETLR) for Databricks. https://discord.gg/p9bzqGybVW
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while control…
Apache Spark - A unified analytics engine for large-scale data processing
An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs
A free to use dbt package for creating and loading Data Vault 2.0 compliant Data Warehouses (powered by dbt, an open source data engineering tool, registered trademark of dbt Labs)
A starter template for Equinor data science / data engineering projects
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Implementation of RelGAN: Relational Generative Adversarial Networks for Text Generation
Open source neural network chess engine with GPU acceleration and broad hardware support.
A logical, reasonably standardized, but flexible project structure for doing and sharing data science work.
Tensorflow implementation of paper http://arxiv.org/abs/1809.02105