You will lead the architecture and implementation of scalable batch and streaming data pipelines using Apache Spark and AWS Glue. Additionally, you will manage ACID-compliant data lakehouse architectures and optimize analytical workloads for high performance.
This remote role is based in Costa Rica and only open to citizens and permanent residents of Costa Rica who do not need visa sponsorship.
We are seeking an experienced Senior Data Engineer to design, build, and optimize our next-generation transactional data lake and analytics infrastructure. In this role, you will lead the architecture and implementation of scalable batch and streaming pipelines using Apache Spark, Apache Hudi, AWS Glue, and AWS Athena, ensuring high performance, transactional consistency, and low-latency querying for enterprise data.
Responsibilities
Data Lake Architecture: Design, implement, and maintain scalable, ACID-compliant data lakehouse architectures utilizing Apache Hudi on AWS S3.
Pipeline Development: Build robust, fault-tolerant ETL/ELT pipelines using Apache Spark (PySpark/Scala) and AWS Glue for high-volume batch and real-time streaming data.
Query Optimization: Write advanced, performant SQL queries and optimize interactive analytical workloads using AWS Athena.
Data Lakehouse Management: Manage Hudi table operations including upserts, soft/hard deletes, compaction, clustering, time-travel queries, and schema evolution.
Performance Tuning: Optimize Glue jobs, Spark cluster performance, partitioning strategies, file formats (Parquet/ORC), and query latency across large-scale datasets.
Governance & Quality: Enforce data quality check frameworks, metadata management, cataloging via AWS Glue Data Catalog, and data security/privacy standards.
Collaboration & Leadership: Mentor junior/mid-level data engineers, participate in code reviews, and partner with Analytics, Data Science, and Product teams to align data modeling with business goals.
Technical Skills
Experience: 5+ years of software/data engineering experience, with a track record of architecting distributed data platforms.
SQL: Advanced mastery of complex SQL queries, analytical functions, CTEs, and performance tuning techniques.
Apache Spark: Deep technical knowledge of Spark architecture (RDDs, DataFrames, Spark SQL, Structured Streaming), memory management, and job optimization.
Apache Hudi: Hands-on experience implementing transactional data lakes with Apache Hudi (Copy-on-Write vs. Merge-on-Read, indexing strategies, CDC ingestion, compaction).
AWS Glue: Expertise in building serverless Glue ETL pipelines, custom scripts, dynamic frames, Glue Triggers, and managing the AWS Glue Data Catalog.
AWS Athena: Strong proficiency with ad-hoc querying, query execution cost/time optimization, dynamic partitioning, and federated queries.
Programming: Proficiency in Python or Scala.
Data Warehousing & Modeling: Solid grasp of dimensional modeling, data vault, Star/Snowflake schemas, and database internals.
Nice to Have
Experience setting up CDC (Change Data Capture) pipelines (e.g., AWS DMS, Debezium) into Apache Hudi
Experience with Infrastructure as Code (Terraform, AWS CloudFormation, or AWS CDK)
Knowledge of orchestration tools such as Apache Airflow or AWS Step Functions.
AWS Certified Data Engineer – Associate or AWS Certified Big Data / Data Analytics Specialty
Strategic Skills
Excellent verbal and written communication skills.
Team player.
Experience working within agile environments.
This remote role is based in Costa Rica and only open to citizens and permanent residents of Costa Rica who do not need visa sponsorship.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”