Stars
The world's fastest open query engine for sub-second analytics both on and off the data lakehouse. With the flexibility to support nearly any scenario, StarRocks provides best-in-class performance …
Apache XTable (incubating) is a cross-table converter for lakehouse table formats that facilitates interoperability across data processing systems and query engines.
Apache Polaris, the interoperable, open source catalog for Apache Iceberg
Remote shuffle service for Apache Spark to store shuffle data on remote servers.
Apache Flink 源码分析系列,基于 git tag 1.1.2
大数据相关内容汇总,包括分布式存储引擎、分布式计算引擎、数仓建设等。关键词:Hadoop、HBase、ES、Kudu、Hive、Presto、Spark、Flink、Kylin、ClickHouse
Nessie: Transactional Catalog for Data Lakes with Git-like semantics
Apache Druid: a high performance real-time analytics database.
A site for knowledge sharing and neat things.
An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
A software library of stochastic streaming algorithms, a.k.a. sketches.
Java library for the HyperLogLog algorithm
Mirror of the official PostgreSQL GIT repository. Note that this is just a *mirror* - we don't work with pull requests on github. To contribute, please see https://wiki.postgresql.org/wiki/Submitti…
Sampling CPU and HEAP profiler for Java featuring AsyncGetCallTrace + perf_events
Distributed reliable key-value store for the most critical data of a distributed system
A Kubernetes toolkit for building distributed applications using cloud native principles
Algorithm Patterns — the most scientific way to practice, the fastest path to an offer. You deserve it~ 算法模板,最科学的刷题方式,最快速的刷题路径,你值得拥有~
ClickHouse® is a real-time analytics database management system
CockroachDB — the cloud native, distributed SQL database designed for high availability, effortless scale, and control over data placement.
A course to build distributed key-value service based on TiKV model
open source training courses about distributed database and distributed systems