Starred repositories
The context API to search, scrape, and interact with the web at scale. 🔥
Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, a…
New and extensible file format for storage of large columnar datasets.
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux…
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
ClickHouse® is a real-time analytics database management system
Mirror of the official PostgreSQL GIT repository. Note that this is just a *mirror* - we don't work with pull requests on github. To contribute, please see https://wiki.postgresql.org/wiki/Submitti…
Apache DataFusion Ballista Distributed Query Engine
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, an…
Repo for everything open table formats (Iceberg, Hudi, Delta Lake) and the overall Lakehouse architecture
Spark integrations for working with Lance datasets
Apache XTable (incubating) is a cross-table converter for lakehouse table formats that facilitates interoperability across data processing systems and query engines.
LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next…
An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs
A composable and fully extensible C++ execution engine library for data management systems.
A cross platform way to express data transformation, relational algebra, standardized record expression and plans.
NVIDIA cuDF for Apache Spark plugin - accelerate Apache Spark with GPUs
watsonx.data Spark adapter for Data Build Tool (DBT)
Apache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering excellence, developer experience, and…
The RasQberry project: Exploring Quantum Computing and Qiskit with a Raspberry Pi and a 3D Printer
Gluten is a middle layer responsible for offloading JVM-based SQL engines' execution to native engines.
Official implementation of CVPR2020 paper "VIBE: Video Inference for Human Body Pose and Shape Estimation"