Match your resume skills with our AI powered skill match!
You will architect and scale distributed web scraping systems to collect data from thousands of internet sources while ensuring reliable ingestion into a production-grade data product. Additionally, you will maintain the data access layer using GraphQL and collaborate with cross-functional teams to integrate ML components.
Remote – LATAM
We’re looking for a Senior Data Product Engineer to join an initiative focused on transforming an existing research-driven data collection platform into a scalable, production-grade data product for cyber risk analytics.
The project involves collecting and processing large volumes of data from thousands of internet sources, building resilient web scraping systems, and making the resulting data reliably available to SaaS teams through a structured data access layer.
You’ll work at the intersection of Data Collection, Data Science, and Data Engineering, owning the path from raw web data collection and ingestion through to a reliable, documented data product.
Architect and scale distributed web scraping systems to reliably collect data from thousands of internet sources.
Build strategies to navigate anti-scraping mechanisms such as Cloudflare, CAPTCHAs, rate limiting, IP banning, and browser fingerprinting.
Work with proxy pools, headless browsers, and adaptive crawling techniques to ensure reliable data collection at scale.
Process and extract data from HTML and JSON using tools such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, or CSS selectors.
Refactor and productionize existing Python data collection pipelines, improving reliability, observability, error handling, retries, and alerting.
Build schedulable and containerized ingestion workflows using Airflow or equivalent orchestration tools.
Design PostgreSQL schemas, views, partitioning, and data access patterns for processed web data.
Work with AWS services including S3, EKS, IAM/IRSA, Parameter Store, and ECR.
Build and maintain the data access layer using GraphQL and Hasura.
Collaborate with Data Science and SaaS teams to integrate existing ML components and define reliable API contracts.
Take end-to-end ownership of the data product, contribute to code reviews, and drive technical improvements across the project.
Web Scraping: Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml
Core & Data: Python, PostgreSQL, Airflow, Redis, Elasticsearch
AWS: boto3, S3, EKS, IAM/IRSA, Parameter Store, ECR
Data Access: GraphQL, Hasura
Infrastructure & DevOps: Docker, Kubernetes, Helm, GitHub Actions
Additional technologies: Liquibase/Flyway, SQLAlchemy
Strong hands-on experience with advanced web scraping at scale.
Experience overcoming anti-scraping and bot-protection mechanisms, including proxy rotation, headless browsers, rate limiting, IP blocking, fingerprinting, or similar challenges.
Strong experience with web scraping and DOM parsing tools such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, or CSS selectors.
Strong professional experience with Python and data processing.
Hands-on experience with AWS, particularly S3 and boto3; experience with EKS, IAM/IRSA, Parameter Store, or ECR is highly valuable.
Experience building reliable data pipelines using Airflow or an equivalent orchestrator.
Strong knowledge of PostgreSQL and relational database fundamentals.
Ability to take end-to-end ownership of technical solutions and work autonomously.
Experience contributing to code reviews and maintaining high software engineering standards.
Experience with GraphQL, Hasura, Redis, Elasticsearch, Docker, or Kubernetes is a plus.
Experience with cybersecurity data, NLP/ML pipelines, or data-as-a-product environments is a plus.
English level B2 or higher, with the ability to communicate effectively with technical and cross-functional teams.
Time off & well-being: vacations fully flexible and self-managed, sick leave and personal days, public holidays, paternity and maternity leave, study leave, and moving days.
Learning & growth: training in best practices and tech, books and light talks, in-house English classes, continuous feedback, and 1:1 career development sessions.
Work experience: flexible working hours, equipment and work materials provided, internal events and team activities, and a day off on your birthday.
Contract & setup: 100% remote positions across LATAM, under a contractor model with payment in USD.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Development
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”