Designing Reliable CDC Pipelines on AWS
A practical architecture for handling full loads, updates, deletes, late-arriving changes, and replay without turning the pipeline into a collection of special cases.
I design and write about reliable data systems—distributed processing, change data capture, data quality, cloud architecture, open-source technologies, and transportation safety analytics.
A practical architecture for handling full loads, updates, deletes, late-arriving changes, and replay without turning the pipeline into a collection of special cases.
How freshness, historical baselines, null rates, duplicates, and distribution checks can catch failures that a simple row-count test will miss.
Nested field statistics and inferred metrics limits in the Python implementation of Apache Iceberg.
Regression fix for optional defaults in the PostgreSQL replication source.
I am a data engineer focused on reliable cloud data systems, distributed processing, and trustworthy analytics. I use this site to publish technical essays, architecture lessons, open-source work, and ideas at the intersection of data engineering, transportation, and public safety.