Build and maintain AWS-based data platforms, pipelines, storage, APIs, and integrations for scientific, geospatial, environmental, and Earth observation data. Implement governance, automation, testing, and observability, and collaborate with scientists and the Senior Architect to deliver reliable, production-ready solutions.
The Data Platform Engineer IV implements and maintains cloud-native AWS data platforms for scientific and environmental workloads. You will build the pipelines, storage, integrations, APIs, automation, governance, and monitoring that move scientific, geospatial, and Earth observation data from source to the researchers and programs that rely on it, working from the Senior Architect's designs while writing clean, tested code and providing feedback on feasibility and trade-offs.
Responsibilities include:
Build cloud-native data platforms on AWS that support scientific and environmental data workloads
Build and maintain ingestion and processing pipelines for large, complex scientific datasets, including Earth observation, weather, and environmental data
Implement storage layers designed for scientific data formats and access patterns (e.g., NetCDF, HDF5, Zarr, GeoTIFF)
Implement data governance, metadata management, and lifecycle controls so federal and scientific data is trusted, discoverable, and well managed, with attention to provenance and reproducibility
Develop integrations with internal systems, external scientific data sources, and third-party services
Build APIs and services that expose data and platform capabilities to scientists, analysts, and downstream systems
Work with scientists and data users to understand how data is produced and consumed, and make sure the platform serves their needs
Implement cloud-based data platform components from solution designs and direction provided by the Senior Architect
Automate deployment, orchestration, testing, and operations using infrastructure as code and CI/CD practices
Add monitoring, logging, and alerting so pipelines are observable and reliable
Give the Senior Architect feedback on feasibility and trade-offs during implementation
Write clean, tested, well-documented code and participate in code reviews
Perform other duties and responsibilities as assigned
What You Will Bring
Basic Qualifications
Bachelor's degree in Computer Science, Engineering, a physical, environmental, or Earth science, or a related field, or equivalent practical experience
8+ years of professional software or data engineering experience, including building production data platforms
Demonstrated experience building pipelines or platforms for scientific, geospatial, environmental, or Earth observation data
Hands-on experience with scientific or geospatial data formats (e.g., NetCDF, HDF5, Zarr, GeoTIFF)
Strong proficiency in Python and SQL
Hands-on experience implementing solutions in a major cloud environment, with AWS strongly preferred
Experience building data ingestion and processing pipelines with workflow or orchestration tools
Experience with cloud data storage technologies (object storage, data lakes, warehouses, databases)
Experience implementing data governance, cataloging, and lifecycle management practices
Experience delivering solutions in federal or other regulated environments, including security and compliance requirements (e.g., NIST, FedRAMP)
Experience building APIs and system integrations
Experience with infrastructure as code and CI/CD tooling
Solid software engineering fundamentals: testing, version control (Git), code review, and documentation
Ability to interpret architectural designs and implement them as scalable, production-ready solutions
Strong communication skills and the ability to collaborate with architects, scientists, and cross-functional teams
Preferred Qualifications
Experience supporting Earth observation, weather, climate, or environmental data platforms
Experience working in scientific research, federal science, or government environments
Experience with distributed processing frameworks (e.g., Spark, Dask, Flink) and streaming technologies (e.g., Kafka, Kinesis)
Familiarity with modern data and table formats (e.g., Parquet, Avro, Iceberg, Delta Lake)
Experience with data governance, metadata management, and data catalog tools
Experience with monitoring and observability tools (e.g., CloudWatch, Prometheus, Grafana, Datadog)
Background in cloud security, IAM, and cost optimization (FinOps)
Experience with containerization (Docker)
AWS certification (e.g., Solutions Architect, Data Engineer, or DevOps Engineer)
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”