Data Engineer (Apache Spark)

 Posted an hour ago
  
 Worldwide
  
5-10 years experience
Apply Now

Please mention DailyRemote when applying

AI Summary

Design and operate high-volume analytical data systems end to end using Apache Spark. Build and maintain high-scale batch and near-real-time data pipelines on self-managed infrastructure.

This is a remote position.

Building high-scale batch and near-real-time data pipelines deployed on infrastructure we run ourselves (on-prem), not managed cloud services. You will design and operate high-volume analytical data systems end to end, with Apache Spark as the core processing engine for both batch and streaming workloads.



Requirements


       7+ years of experience in data engineering and software development

       Ability to write high-quality code in Java/Scala, Python, or equivalent languages

       Deep, hands-on production experience with Apache Spark — batch and Spark Structured Streaming (core requirement)

       Demonstrated Spark performance tuning: partitioning, caching and persistence, broadcast joins, shuffle reduction, data-skew handling, and Adaptive Query Execution

       Experience operating Spark on self-managed clusters (YARN, Kubernetes, or standalone) — executor sizing, resource allocation, and multi-tenant workloads

       Practical experience with Kafka (or equivalent messaging systems) as a Spark source and sink for high-volume workloads, including offset and checkpoint management

       Practical experience with distributed query engines (e.g., Trino/Presto or similar)

       Practical experience with ETL / data integration tools, commercial or open-source (e.g., Datastage, Informatica, Apache NiFi, or similar)

       Practical experience with SQL-based transformation frameworks (e.g., dbt or others)

       Strong SQL skills and understanding of data modeling and data warehousing for analytical workloads

       Hands-on experience with real-time / low-latency analytical stores (columnar or OLAP engines, e.g., Apache Pinot/ClickHouse or similar)

       Practical experience with big-data platforms and distributions (e.g., Cloudera, Hadoop ecosystem, Databricks, or similar)

       Practical experience containerizing and operating data workloads (Docker; Kubernetes a plus)

       Experience with workflow orchestration tools (e.g., Airflow or similar)

       Familiarity with data lake table formats (e.g., Apache Iceberg, Delta Lake, or similar), including schema evolution and compaction

       Familiarity with data governance / cataloging tools (e.g., DataHub or similar)

       Familiarity with lakehouse management systems (e.g., Apache Amoro or similar)

       Familiarity using AI tools for development and debugging (Claude, Cursor, Codex)



Similar Jobs

See all Remote Software Development jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Data Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified