For Employers

RYZ Labs

Data Engineer

Posted a day ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review

Questions interviewers often ask for this role, with sample answers.

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

Own and improve batch and event-driven data ingestion pipelines, including data cleaning, entity resolution, validation, and publishing across AWS-based systems. Monitor data quality and pipeline reliability, troubleshoot production incidents, and collaborate with engineering, product, and analytics teams to ensure data reaches downstream services correctly.

Remote position - only for professionals based in LATAM.

 

 

We are looking for a Senior Data Engineer for one of our client's teams. You will take ownership of the systems responsible for ingesting, transforming, validating, and publishing data across our platform.

 

In this role, you will work at the data ingestion boundary, where data from multiple external sources enters our systems. You will be responsible for building reliable pipelines, resolving identity and entity-matching challenges, detecting data-quality issues, and ensuring data flows correctly into downstream products and services.

 

This is a hands-on engineering role that combines AWS data engineering, Python/SQL development, event-driven architectures, data quality, and identity resolution. You will collaborate closely with engineering, product, analytics, and downstream platform teams to troubleshoot issues and improve the reliability of our data ecosystem.

 

 

What You'll Do

• Own and evolve data ingestion pipelines, including source ingestion/scraping, data cleaning and curation using AWS Glue, entity resolution, and event-driven data publishing.

• Design and maintain reliable batch and event-driven data workflows across our data platform.

• Diagnose and resolve identity-resolution issues across multiple data sources, including deduplication and entity-matching challenges.

• Develop and maintain data quality and validation frameworks, including freshness monitoring, null-rate checks, consistency validation, and schema-drift detection.

• Monitor data at the ingestion boundary and proactively identify issues before they impact downstream systems and products.

• Partner with engineering teams responsible for the events bridge and downstream identity services to trace data and events end-to-end.

• Collaborate with Product, Assessments, Analytics, and other stakeholders to understand how data flows into downstream products and ensure those requirements are reflected in the data pipelines.

• Automate infrastructure and pipeline changes using Infrastructure as Code, primarily Terraform or AWS CDK.

• Work with AWS services including Glue, Athena, S3, and event-driven services to build and operate scalable data infrastructure.

• Participate in production incident response, quickly identifying the scope and impact of data-quality issues and implementing remediation.

• Continuously improve pipeline reliability, observability, maintainability, and operational processes.

 

What You'll Bring

• 5+ years of experience building and operating production-grade data pipelines.

• Strong experience working with both batch data processing and event-driven architectures.

• Hands-on experience with AWS Glue, Athena, and S3-based data lakes, including layered/medallion-style data transformations.

• Strong proficiency in Python and SQL for data transformation, processing, and analysis.

• Experience with identity resolution, entity matching, or deduplication, including exact and fuzzy matching approaches.

• Understanding of the challenges and tradeoffs involved in maintaining durable identifiers and first-seen/locked identity mappings.

• Experience working with event schemas and schema-registry-backed contracts, such as Protobuf, and an understanding of the risks associated with schema changes.

• Strong troubleshooting and production incident-response skills, including the ability to determine impact, identify root causes, and implement fixes quickly.

• Ability to understand and debug code written in a functional or concurrent programming language, such as Elixir.

• Strong communication and collaboration skills, with the ability to work effectively across engineering, product, analytics, and other technical teams.

 

Nice to Have

• Experience with Elixir/Phoenix or another BEAM-based concurrent processing framework.

• Experience with DynamoDB-backed identity, lookup, or matching services.

• Experience with the Snowflake ecosystem, including data modeling, Snowpipe, Streams, and Tasks.

• Experience working with sports data providers, such as MLB.com or Sportradar, or with other licensed-content/data provider ecosystems.

• Experience implementing data observability, including freshness/staleness alerts, null-rate monitoring, schema-drift detection, and data-quality dashboards.

• Experience working with Infrastructure as Code using Terraform or AWS CDK.

 

Technologies

Languages: Python, SQL, familiarity with Elixir

AWS: Glue, Athena, S3, DynamoDB

Data & Streaming: Data Lakes, Event-Driven Architecture, Protobuf, Schema Registries

Infrastructure: Terraform, AWS CDK

Data Platforms: Snowflake

Engineering Practices: Data Quality, Data Observability, Identity Resolution, Entity Matching, Incident Response

\n


\n

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Remote Satellite Systems I&T Engineer

Freelance United States $50 - $70 per hour Software Development

Network Cloud Security Specialist

Full Time Germany Software Development

Solution Architect AI & ML for Life Sciences

Full Time United States $119K - $221K per year Software Development

Salesforce QA Analyst - LATAM

Full Time, Freelance Brazil Software Development

Principal Marketing Analytics Consultant

Full Time United States $70000 - $110K per year Software Development

Full Stack Developer

Full Time United States $110K - $114K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 217,169+ Jobs in Data Engineer

Answer easy questions

Answer easy questions

217,169+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified