- Beijing, China
- https://carp84.github.io/
- @LiyuApache
- in/carp84
Lists (13)
Sort Name ascending (A-Z)
CICD
Tools and frameworks for CICDCloud-native (DisAgg)
Dis-aggregation related techDatalake
Datalake related techDB&DW&Lakehouse
OLTP/OLAP/HTAP/HSAP/StreamingWarehouse/Lakehouse, etc.ETL & ELT
Data integration and transform projects.Federation and Private Computing
Repositories about federation computing with privacy and security considerationFlink and related
Flink related projectsJVM and Java related
JVM and Java relatedStars
Apache Ossie, industry wide specification effort to standardize how we exchange semantic metadata across analytics, AI and BI platforms, providing a vendor neutral, single source of truth for seman…
Skills for Real Engineers. Straight from my .agents directory.
Lightweight coding agent that runs in your terminal
An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of…
A modular framework for evaluating disk-based vector search, enabling controlled ablations and end-to-end comparisons across diverse storage environments, concurrency regimes, and query distributions.
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
Self-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.
Scalable data pre processing and curation toolkit for LLMs
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
open-source agentic AI data assistant for the next generation of AI + Data products.
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux…
Supercharge Your LLM Application Evaluations 🚀
Flink Agents is an Agentic AI framework based on Apache Flink
Model Context Protocol Servers for Milvus
DuckLake is an integrated data lake and catalog format
A lightweight data processing framework built on DuckDB and 3FS.
A high-performance distributed file system designed to address the challenges of AI training and inference workloads.
Apache Fluss is a streaming storage built for real-time analytics.
Benchmarks of approximate nearest neighbor libraries in Python
DuckDB is an analytical in-process SQL database management system
World's most powerful open data catalog for building a high-performance, geo-distributed and federated metadata lake.
Open, Multi-modal Catalog for Data & AI
Apache Polaris, the interoperable, open source catalog for Apache Iceberg
A collection of modern C++ libraries, include coro_http, coro_rpc, compile-time reflection, struct_pack, struct_json, struct_xml, struct_pb, easylog, async_simple etc.
Apache DataFusion Comet Spark Accelerator