See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewQuestions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
Design and operate high-volume analytical data systems end to end using Apache Spark. Build and maintain high-scale batch and near-real-time data pipelines on self-managed infrastructure.
This is a remote position.
Building high-scale batch and near-real-time data pipelines deployed on infrastructure we run ourselves (on-prem), not managed cloud services. You will design and operate high-volume analytical data systems end to end, with Apache Spark as the core processing engine for both batch and streaming workloads.
• 7+ years of experience in data engineering and software development
• Ability to write high-quality code in Java/Scala, Python, or equivalent languages
• Deep, hands-on production experience with Apache Spark — batch and Spark Structured Streaming (core requirement)
• Demonstrated Spark performance tuning: partitioning, caching and persistence, broadcast joins, shuffle reduction, data-skew handling, and Adaptive Query Execution
• Experience operating Spark on self-managed clusters (YARN, Kubernetes, or standalone) — executor sizing, resource allocation, and multi-tenant workloads
• Practical experience with Kafka (or equivalent messaging systems) as a Spark source and sink for high-volume workloads, including offset and checkpoint management
• Practical experience with distributed query engines (e.g., Trino/Presto or similar)
• Practical experience with ETL / data integration tools, commercial or open-source (e.g., Datastage, Informatica, Apache NiFi, or similar)
• Practical experience with SQL-based transformation frameworks (e.g., dbt or others)
• Strong SQL skills and understanding of data modeling and data warehousing for analytical workloads
• Hands-on experience with real-time / low-latency analytical stores (columnar or OLAP engines, e.g., Apache Pinot/ClickHouse or similar)
• Practical experience with big-data platforms and distributions (e.g., Cloudera, Hadoop ecosystem, Databricks, or similar)
• Practical experience containerizing and operating data workloads (Docker; Kubernetes a plus)
• Experience with workflow orchestration tools (e.g., Airflow or similar)
• Familiarity with data lake table formats (e.g., Apache Iceberg, Delta Lake, or similar), including schema evolution and compaction
• Familiarity with data governance / cataloging tools (e.g., DataHub or similar)
• Familiarity with lakehouse management systems (e.g., Apache Amoro or similar)
• Familiarity using AI tools for development and debugging (Claude, Cursor, Codex)
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 219,827+ Jobs in Data Engineer
Answer easy questions
219,827+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”