Data Engineer | Apache Spark, Flink, Kubernetes, Trino, dbt | Building Open-Source Data Lakehouses
I'm a Data Engineer specializing in building open-source, production-grade data lakehouses on Kubernetes. Over the past year and a half, I designed and operated a complete self-hosted data platform — processing ~5 million rows through a full Bronze, Silver, and Gold medallion architecture using Apache Spark, Apache Flink, Trino, dbt-Trino, Apache Iceberg, and Google BigQuery. My core stack includes Python, SQL, and PySpark for data processing; dbt and Trino for transformation; Dagster for orchestration; and a full CI/CD pipeline with GitHub Actions and Mage AI for automated deployments to Kubernetes. I also built a real-time streaming pipeline using Apache Flink and Kafka-compatible messaging with Kappa architecture, sinking data into Snowflake. Key achievements include implementing a layered data quality framework (PyDeequ, dbt tests, Elementary), building 8 BI dashboards in Apache Superset, setting up full observability with Grafana and Prometheus, and resolving 40+ real production issues — from Kubernetes memory crashes to cloud IAM permission chains. I hold HackerRank certifications in SQL (Advanced) and completed Spark Fundamentals through Cognitive Class. My BSc in Computer Science (GPA 3.33/4.0) gave me the foundation; everything else in modern data engineering I taught myself by building real systems. I'm looking for a remote data engineering role where I can bring this end-to-end pipeline ownership — from raw ingestion to production dashboards — to a team solving real business problems. I'm comfortable working async, adapting to different time zones, and taking ownership of complex technical challenges independently.
Member Since
July 17, 2026
Last Active
15 minutes ago