I'm a data engineer, building real-time streaming systems that power machine learning and AI. Over the past decade I've worked extensively with Kafka, Flink, and Spark, designing high-throughput pipelines and lakehouse architectures that stay resilient, observable, and maintainable at scale.
What excites me most is where streaming meets intelligence: feeding live data into models for online learning, real-time detection, and adaptive decision-making. I'm increasingly focused on applying ML and AI to streaming data, turning fast-moving events into systems that learn and respond in the moment.
As a passionate engineer and writer, I share practical insights on real-time analytics, data architectures, data lakehouse patterns, and data lineage. I recently presented "Building End-to-End Data Lineage with Kafka, Flink, and Spark" at Current and Programmable in 2026, and I write regularly on my blog.
Blog: https://jaehyeon.me
Open source:
- dynamic-des (dynamic discrete event simulation & digital twins)
- odctl (Open Data Stack)