Apply Now

Please mention DailyRemote when applying

AI Summary

Design, build, and operate real-time streaming pipelines and analytical data platforms using Apache Kafka, Flink, and cloud-native technologies. Manage infrastructure as code, maintain data lakehouse architectures, and support cloud migration initiatives while ensuring high system reliability.

Company Background

Our client is a leading global data center provider delivering hyperscale and edge infrastructure solutions across the Americas, EMEA, and Asia-Pacific. With 80+ data centers in 20+ countries, they partner with industry leaders such as Google, Oracle, NVIDIA, and Microsoft Azure to power the world’s digital infrastructure. Recognized as a USA TODAY Top Workplace for four consecutive years, the company continues to expand its global footprint and customer ecosystem.

Project Description

The project focuses on building a cloud-based data platform for processing high-volume sensor and telemetry data. The specialist will help develop and operate real-time and analytical data capabilities that support dashboards, applications, and downstream data consumers, while contributing to the platform’s migration and evolution in AWS.

Technologies

  • Apache Kafka, Apache Flink
  • Java, Python, SQL
  • Apache Iceberg / Delta Lake / Hudi
  • Parquet, Avro, Schema Registry
  • Amazon Athena / Spark SQL
  • ClickHouse / Druid / Pinot
  • SQL Server / PostgreSQL
  • Kubernetes (Amazon EKS)
  • Helm, Argo CD, GitOps
  • Terraform, GitHub Actions
  • AWS (Amazon MSK, EMR Serverless, MWAA, S3, Athena, IAM)
  • LLM, Embeddings, Vector Search

What You'll Do

  • Design, build, and operate real-time streaming pipelines with Apache Kafka (Amazon MSK) and Apache Flink for high-throughput sensor and telemetry data;
  • Define and manage streaming data contracts, including Avro schemas and schema evolution through a schema registry;
  • Build and maintain analytical serving layers using ClickHouse or similar columnar OLAP databases, and develop REST APIs for dashboards, applications, and downstream teams;
  • Develop and operate batch and scheduled data workflows with Apache Airflow (Amazon MWAA);
  • Build and operate a data lakehouse based on Apache Iceberg and Amazon S3, using PySpark on EMR Serverless and Amazon Athena;
  • Deploy and operate containerized data workloads on Kubernetes (Amazon EKS) using Argo CD, Helm, and GitOps practices;
  • Manage cloud infrastructure as code with Terraform and support CI/CD automation with GitHub Actions;
  • Monitor production data platforms with Prometheus and Grafana, troubleshoot issues, and participate in incident response;
  • Contribute to cloud migration initiatives by porting data pipelines from existing platforms, including Azure-based data platforms, to AWS;
  • Develop production-grade Java, Python, and SQL code with automated testing;
  • Use AI-assisted development tools to accelerate analysis and implementation while maintaining code quality and architectural integrity;
  • Keep technical documentation and operational runbooks current;

Job Requirements

  • 7+ years of experience in data or software engineering, including production experience with streaming systems;
  • Ability to work independently in ambiguous and fast-changing environments;
  • Deep hands-on experience with Apache Kafka and Apache Flink or an equivalent stream-processing framework, including Java;
  • Strong Python and SQL skills across transactional and analytical databases;
  • Experience with Apache Iceberg, Delta Lake, or Hudi; Parquet, Avro, schema registries, Athena, or Spark SQL;
  • Production experience with ClickHouse, Druid, Pinot, or similar analytical databases, as well as SQL Server or PostgreSQL;
  • Experience with Kubernetes, Amazon EKS, Helm, Argo CD, and GitOps;
  • Experience with Terraform and GitHub Actions;
  • Strong knowledge of AWS services, including MSK, EMR Serverless, MWAA, S3, Athena, and IAM;
  • Knowledge of ML fundamentals, including feature engineering, model training and evaluation, and ML data requirements;
  • Familiarity with LLMs, prompt-based workflows, embeddings, vector search, anomaly detection, and forecasting;
  • Experience with observability, alerting, troubleshooting, and incident response;
  • Ability to communicate technical decisions clearly to technical and non-technical stakeholders;

Nice to Have

  • Working knowledge of Azure Event Hubs, Data Factory, and Synapse;
  • Experience with OT or industrial telemetry, including OPC UA, BMS/EPMS, or time-series sensor data;
  • Experience with cloud-to-cloud or on-premises-to-cloud migrations;
  • Experience with data center or other critical-infrastructure operations;
  • Familiarity with data quality frameworks, data contracts, and metadata management;

What Do We Offer

The global benefits package includes:

  • Technical and non-technical training for professional and personal growth;
  • Internal conferences and meetups to learn from industry experts;
  • Support and mentorship from an experienced employee to help you professional grow and development;
  • Health insurance;
  • Sports activities to promote a healthy lifestyle;
  • Flexible work options, including remote and hybrid opportunities;
  • Referral program for bringing in new talent;
  • Work anniversary program and additional vacation days.

Similar Jobs

See all Remote Software Development jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Data Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified