data load tool (dlt) is an open source Python library that makes data loading easy 🛠️
-
Updated
Sep 21, 2026 - Python
data load tool (dlt) is an open source Python library that makes data loading easy 🛠️
lakeFS - Data version control for your data lake | Git for data
Apache Kyuubi is a distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses.
Few projects related to Data Engineering including Data Modeling, Infrastructure setup on cloud, Data Warehousing and Data Lake development.
BitSail is a distributed high-performance data integration engine which supports batch, streaming and incremental scenarios. BitSail is widely used to synchronize hundreds of trillions of data every day.
An end-to-end GoodReads Data Pipeline for Building Data Lake, Data Warehouse and Analytics Platform.
Apache Iceberg REST Catalog in Rust — access control, credential vending and audit for every engine and AI agent. Apache 2.0.
Apache Amoro(incubating) is a Lakehouse management system built on open data lake formats.
Kylo is a data lake management software platform and framework for enabling scalable enterprise-class data lakes on big data technologies such as Teradata, Apache Spark and/or Hadoop. Kylo is licensed under Apache 2.0. Contributed by Teradata Inc.
Personal Data Engineering Projects
An efficient storage and compute engine for both on-prem and cloud-native data analytics.
Data API Framework for AI Agents and Data Apps
This repository has been merged into Canner/WrenAI under the core/ directory
Enterprise-grade, production-hardened, serverless data lake on AWS
Generic Data Ingestion & Dispersal Library for Hadoop
Real Time Big Data / IoT Machine Learning (Model Training and Inference) with HiveMQ (MQTT), TensorFlow IO and Apache Kafka - no additional data store like S3, HDFS or Spark required
GigAPI is a Timeseries lakehouse for real-time data and sub-second queries, powered by DuckDB OLAP + Parquet Query Engine, Compactor w/ Cloud-Native Storage. Drop-in FDAP alternative ⭐
Use SQL to build ELT pipelines on a data lakehouse.
BtrBlocks: Efficient Columnar Compression for Data Lakes (SIGMOD 2023 Paper)
中国 A 股数据基础设施。42 个日更数据集:行情、基本面、资金面、公告事件、指数行业、宏观与风险。行级溯源、PIT 语义、复权与历史成分内置,MCP 原生。自托管,零注册、零 API Token
To associate your repository with the data-lake topic, visit your repo's landing page and select "manage topics."