Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
Design and build a metadata-driven ingestion framework using reusable pipeline templates on Google Cloud Platform. Manage the end-to-end data lifecycle, including ingestion, transformation, and reconciliation across Raw, Bronze, Silver, and Gold layers.
We are looking for an experienced Data Engineer to design, build, and maintain a scalable, metadata-driven data ingestion and transformation framework on Google Cloud Platform (GCP).
The role will be responsible for developing configuration-driven ingestion pipelines, onboarding multiple source systems, and implementing Raw, Bronze, Silver, and Gold data layers. The Data Engineer will use technologies including BigQuery, Cloud Composer/Apache Airflow, Python, Dataform, Datastream, Pub/Sub, Dataflow, Cloud Run, and Cloud Storage to deliver reliable and reusable data solutions.
The ideal candidate will have strong experience in metadata-driven pipeline development, CDC, advanced SQL, Python, data quality, reconciliation, automated testing, and CI/CD. The engineer will work closely with Data Governance and other platform stakeholders to ensure that data pipelines are scalable, governed, auditable, testable, and production-ready.
Design and build a metadata-driven and configuration-driven ingestion framework using reusable pipeline templates.
Design, develop, and maintain the metadata configuration model covering source systems, source objects, load configurations, execution parameters, run logs, audit information, and reprocessing.
Build dynamic Cloud Composer / Apache Airflow DAGs that generate and execute ingestion workflows based on metadata configurations.
Develop reusable ingestion patterns and pipeline generators instead of building separate pipelines for individual source tables.
Onboard multiple source systems, including relational databases, SAP, REST APIs, files, streaming sources, and other enterprise platforms.
Implement data ingestion into Raw and Bronze layers following defined lakehouse architecture and engineering standards.
Implement source-to-target reconciliation to ensure completeness and accuracy of ingested data.
Implement robust operational controls, including error handling, retries, quarantine processing, alerting, monitoring, and audit logging.
Develop replay-by-batch and reprocessing capabilities to support failed loads, historical reloads, and controlled data recovery.
Support batch, incremental, streaming, and Change Data Capture (CDC) ingestion patterns.
Implement Bronze-layer processing, including data typing, cleansing, schema validation, and schema enforcement.
Implement data de-duplication using appropriate source keys, primary keys, or business keys.
Implement soft-delete representation and appropriate handling of deleted source records.
Implement CDC change-history materialization and maintain historical changes where required.
Develop appropriate BigQuery partitioning, clustering, MERGE, and incremental processing strategies.
Perform data compaction and other performance optimization activities where required.
Design and develop Silver and Gold transformation models using Dataform.
Develop reusable transformation components, macros, dependencies, and incremental transformation strategies.
Implement assertions and automated tests for transformation models and ensure no untested transformation reaches production.
Implement in-pipeline data quality checks, validation rules, and automated promotion gates.
Work closely with the Data Governance Consultant to incorporate data quality, governance, lineage, audit, and control requirements into the engineering framework.
Prevent data that fails critical quality requirements from being promoted to downstream layers.
Implement geospatial data ingestion, including GeoJSON processing and conversion to BigQuery GEOGRAPHY.
Ensure appropriate preservation and handling of Spatial Reference System (SRS) information for geospatial datasets.
Implement ingestion and management of semi-structured and unstructured data using Google Cloud Storage and BigQuery object tables.
Follow Git-based development practices, including branching, pull/merge requests, peer reviews, and code reviews.
Develop automated tests as a standard part of pipeline and transformation development.
Integrate data pipelines and Dataform transformations with CI/CD processes.
Document each source-system onboarding, including configuration, mappings, dependencies, reconciliation, operational procedures, and troubleshooting guidance.
Develop and maintain operational runbooks and onboarding documentation that enable ESNAD teams to independently onboard additional data sources.
Minimum 5+ years of professional Data Engineering experience.
Minimum 2+ years of hands-on experience building data pipelines on Google Cloud Platform (GCP).
Strong hands-on experience with Google Cloud Platform data engineering services.
Expert-level SQL skills, including complex analytical SQL development.
Strong hands-on experience with Google BigQuery, including data modeling and large-volume data processing.
Strong Python programming skills for data engineering, automation, API integration, validation, and framework development.
Strong hands-on experience with Cloud Composer / Apache Airflow.
Proven experience developing dynamic Airflow DAGs.
Experience developing metadata/configuration-driven DAG generation.
Experience with Airflow sensors, custom operators, task dependencies, scheduling, retries, and failure management.
Proven experience designing and implementing metadata-driven or configuration-driven ingestion frameworks.
Experience developing pipeline generators and reusable ingestion templates, rather than only building individual pipelines for individual tables.
Experience designing metadata configurations for source systems, source objects, load parameters, runtime configurations, audit logs, and reprocessing.
Hands-on experience with Google Cloud Datastream or an equivalent log-based CDC technology.
Strong understanding of CDC patterns, including inserts, updates, deletes, soft deletes, incremental loads, and historical change management.
Experience implementing CDC recovery, reconciliation, replay, and reprocessing patterns.
Hands-on experience with Dataform or dbt.
Strong understanding of transformation model structures, macros, reusable logic, assertions, automated tests, dependencies, and incremental strategies.
Experience developing Raw, Bronze, Silver, and Gold data layers or equivalent medallion/lakehouse architecture.
Experience implementing source-to-target data reconciliation and validation.
Experience implementing exception handling, quarantine processes, retry mechanisms, operational monitoring, and alerting.
Working knowledge of Google Cloud Pub/Sub and event-driven or streaming ingestion patterns.
Working knowledge of Dataflow / Apache Beam for data processing.
Working knowledge of Cloud Run and/or Cloud Data Fusion.
Strong understanding of relational database extraction and database-engine-specific CDC constraints.
Working knowledge of SAP data extraction patterns, including ODP, SLT, or certified SAP connectors.
Experience with REST API ingestion, including authentication, pagination, error handling, and incremental extraction.
Experience with file-based data ingestion and file manifest validation.
Experience implementing automated data quality checks, validation rules, and pipeline quality gates.
Strong knowledge of BigQuery partitioning, clustering, MERGE operations, de-duplication, schema enforcement, and incremental processing.
Strong understanding of data pipeline monitoring, auditability, traceability, and operational support.
Strong experience with Git-based development and version control.
Experience working with peer reviews, code reviews, and controlled development workflows.
Experience writing unit, integration, data quality, or pipeline tests as a standard engineering practice.
Experience implementing or working with CI/CD pipelines for data engineering workloads.
Ability to produce high-quality technical documentation, source onboarding documentation, and operational runbooks.
Google Cloud Professional Data Engineer certification.
Google Cloud Associate Cloud Engineer certification.
dbt Analytics Engineering Certification.
Advanced hands-on experience with Google Cloud Datastream and enterprise-scale CDC implementations.
Advanced experience with Dataflow / Apache Beam for batch and streaming workloads.
Advanced experience with Pub/Sub and event-driven data architectures.
Advanced BigQuery performance optimization and cost-optimization experience.
Experience designing and implementing enterprise-scale lakehouse or medallion architectures.
Experience designing reusable enterprise data ingestion frameworks and platform accelerators.
Advanced hands-on experience with Dataform, including reusable components, assertions, incremental models, and CI/CD.
Experience designing enterprise-level automated data quality frameworks.
Experience with data governance, metadata management, data lineage, data cataloging, auditability, and traceability.
Strong hands-on experience with SAP ODP, SAP SLT, or certified SAP extraction connectors.
Experience implementing SAP CDC and incremental extraction patterns.
Experience processing GeoJSON and BigQuery GEOGRAPHY data.
Knowledge of geospatial data engineering and Spatial Reference Systems (SRS).
Experience handling semi-structured and unstructured data at enterprise scale.
Experience with Google Cloud Storage and BigQuery object tables.
Experience implementing CI/CD pipelines for Dataform, dbt, Airflow, and other data engineering workloads.
Experience with automated deployments, environment promotion, testing gates, and release-management practices.
Knowledge of cloud security, IAM, service accounts, secrets management, and secure data pipeline design.
Experience implementing monitoring, observability, logging, alerting, and operational dashboards for enterprise data pipelines.
Experience working in large enterprise data transformation or cloud migration programs.
Experience working closely with Data Governance, Data Architecture, Security, DevOps, and Business teams.
Experience mentoring junior and mid-level Data Engineers and establishing engineering standards.
For a Senior Data Engineer / Technical Lead, 8+ years of overall Data Engineering experience is preferred, along with demonstrated technical leadership and architecture ownership.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 216,636+ Jobs in Data Engineer
Answer easy questions
216,636+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”