Own and modernize the company's data platform by transitioning to a lakehouse architecture using AWS services and dbt. Manage the deployment, scaling, and orchestration of the data stack while ensuring reliability, security, and cost-efficiency.
Senior Data Engineer
We are looking for a hands-on Senior Data Platform Engineer to own, scale, and modernize the company's data platform, ensuring it is scalable, reliable, secure, cost-efficient, and capable of supporting both operational and analytical workloads.. You will maintain our current platform while actively helping us transition toward a lakehouse architecture: keeping raw data in S3, utilizing Redshift external schemas, and leveraging dbt for optimized transformations.
Key Responsibilities
- Platform & Infrastructure Management: Own the deployment, scaling, and orchestration of our core data stack using AWS Managed Workflows for Apache Airflow (MWAA). Design and evolve the overall data platform architecture, making thoughtful tradeoffs around scalability, reliability, cost, and maintainability.
- Lakehouse Engineering: Configure and optimize raw data storage in S3, managing Redshift external schemas (Spectrum) and AWS Glue Data Catalog. Define efficient partitioning strategies, leverage columnar file formats (Parquet), and optimize storage layouts for query performance, scalability, and cost.
- Data Architecture & Modeling (lakehouse, dimensional models, Redshift). Design scalable analytical data models using dimensional modeling and lakehouse best practices to support reporting, analytics, and downstream consumers.
- Data Transformation: Manage and optimize dbt to ensure all raw data from external schemas is efficiently transformed, tested, and materialized strictly as processed data within Amazon Redshift.
- Data Reliability & Operations (monitoring, testing, incident response). Build observability into the data platform through monitoring, logging, alerting, and operational dashboards.
- Performance & Cost Optimization (Redshift tuning, S3 layout, workloads)
- Governance & Collaboration (security, documentation, working with stakeholders)
- Modernization Pipeline: Help design and implement the transition toward event-based ingestion into S3.
- CI/CD & DevOps: Implement and maintain robust CI/CD deployment pipelines for our dbt projects, Airflow DAGs, and infrastructure.
- BI Support: Ensure high performance, access control, and uptime for Metabase connecting to Redshift.
Technical Skills & Requirements
- Core AWS Stack: Extensive hands-on experience deploying and managing AWS data services, specifically MWAA (Airflow), Redshift / Redshift Spectrum, S3, IAM, and Glue.
- Data Transformation: Advanced proficiency with dbt (structuring dbt projects, configuring sources, writing custom macros, and optimizing incremental models).
- Infrastructure as Code (IaC): Solid hands-on experience deploying AWS data platform components using Terraform or AWS CloudFormation.
- SQL & Performance Tuning: Expert-level SQL skills, with a deep understanding of Redshift distribution/sort keys and optimizing queries across external schemas.
- Programming: Strong Python experience for developing Airflow DAGs, automation, integrations, and data engineering tooling
- Data Storage & Lakehouse: Experience designing efficient data lakes using Parquet, partitioning strategies, metadata catalogs, and external table technologies such as Redshift Spectrum..
- Modern Lakehouse Technologies: Experience with open table formats such as Apache Iceberg, Delta Lake, or Apache Hudi, including an understanding of ACID transactions, schema evolution, time travel, and metadata management.
Bonus / Nice-to-Have
- Experience with Apache Kafka or AWS MSK for event-driven data streaming and real-time ingestion into S3.
- Streaming & Event-Driven Architecture: Experience designing resilient streaming pipelines with appropriate delivery guarantees, reconciliation, backfill strategies, and schema contract management.
- Familiarity with modern analytical query engines such as DuckDB or MotherDuck is a plus.