Define the architecture and migration strategy for modernizing legacy ETL and ELT pipelines to Databricks on AWS. Oversee the decommissioning of legacy DataStage components while ensuring data quality, security, and operational readiness.
This is a remote position.
Onsite/Hybrid/Remote: Remote
Duration: 12 months
Rate Range: $75 in W2
Work Authorization: GC and US Citizens Only
Must Have:
ETL/ELT architecture and modernization
IBM DataStage
Databricks and Apache Spark
Delta Lake, Unity Catalog, and Photon
AWS Glue, Redshift, and Lambda
Python and Unix/Linux scripting
Terraform and CI/CD
Large-scale database and data warehouse migrations
Data governance, quality, and observability
Responsibilities:
Define the architecture and migration strategy for modernizing legacy ETL and ELT pipelines.
Assess IBM DataStage jobs, databases, and data warehouses for migration readiness.
Design scalable data solutions using Databricks Spark on AWS.
Establish architecture standards, reusable frameworks, and governance controls.
Design CI/CD pipelines for automated builds, testing, and deployment.
Lead technical design reviews, migration planning, and artifact validation.
Define parity, functional, UAT, regression, and performance testing strategies.
Ensure schema validation, data quality, lineage, security, and production readiness.
Guide cutover, go-live, hypercare, and operational stabilization activities.
Oversee the decommissioning of legacy DataStage jobs and related components.
Create operational documentation and conduct knowledge-transfer sessions.
Provide technical direction to developers and engineering teams.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”