For Employers

Oxford Data Plan

Senior Data Engineer (Web Scraping)

Posted 4 days ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

You will own the development and operation of scalable web-scraping and data ingestion pipelines within a lakehouse architecture. This involves designing robust frameworks for extraction, monitoring, and failure handling while ensuring compliance with legal and privacy standards.

We are looking for an experienced, self-driven Data Engineer to take ownership of web scraping and web-based data acquisition within our data platform.

We already use web scraping across a number of ingestion processes, and we are looking to improve the maturity, scale and robustness of this capability. This includes integrating scraping workloads into our lakehouse architecture, improving scheduling, monitoring and storage, and establishing reusable patterns for building and operating scrapers reliably.

This is a senior, hands-on role. We are looking for someone who has previously built and operated production web-scraping systems at scale and can independently take a new data source from investigation through implementation, deployment and ongoing support.

Roles and Responsibilities

· Own the development and operation of web-scraping and web-data ingestion pipelines.

· Design and establish a scalable web-scraping framework, including common patterns for extraction, scheduling, storage, monitoring, validation and failure handling.

· Design, build and maintain reliable production scrapers for new and existing data sources.

· Evaluate and make build-versus-buy recommendations for scraping infrastructure and third-party services, considering capability, reliability, cost, operational complexity and risk.

· Investigate websites and determine the most appropriate acquisition approach, including APIs, direct HTTP requests, HTML parsing, browser automation or third-party tooling.

· Ensure scraping activities appropriately consider internal policies, website terms, robots.txt, access restrictions, privacy and intellectual-property constraints, escalating unclear cases where needed.

· Diagnose and resolve issues caused by changing websites, dynamic content, authentication, rate limits and other operational challenges.

· Use AI tooling to work more effectively while understanding, reviewing and being able to defend the code you ship.

Required Qualifications

· Demonstrated professional experience building and operating production web-scraping systems at scale.

· Proven ability to independently own substantial scraping projects from initial investigation through production operation.

· Strong production Python engineering skills, including building maintainable applications rather than standalone scripts.

· Strong practical experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright or Selenium.

· Good understanding of HTTP, HTML, APIs, JavaScript-rendered websites and browser/network behaviour.

· Experience handling common scraping challenges such as pagination, authentication, sessions, retries, rate limiting, concurrency and proxies.

· Strong understanding of data pipelines, data quality and how collected data should be validated, stored and consumed downstream.

· Experience deploying, monitoring and supporting production workloads in a cloud environment.

· Strong debugging and problem-solving skills, with the ability to work independently and make sensible engineering decisions.

· Experience working effectively within a remote engineering team, including code review, documentation and ticket-based workflows.

Desirable Skills

· Experience with AWS.

· Experience with lakehouse or data-lake architectures, particularly Iceberg.

· Experience with PySpark or other distributed data-processing technologies.

· Experience with Docker and containerised workloads.

· Experience with Terraform or other infrastructure-as-code tooling.

· Experience with Grafana or similar observability platforms.

· Experience operating high-volume or distributed crawling systems.

· Experience evaluating or operating commercial scraping, proxy or browser-infrastructure services.

· Experience implementing automated scraper testing, canary runs or source-drift detection.

· Experience working with legal, compliance, privacy or data-governance teams on web-data acquisition.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Freelance Full-Stack Developer - Python & JavaScript

Full Time Philippines Software Development

Freelance Email Developer - Adobe Journey Optimizer (AJO)

Full Time Malaysia Software Development

Sr Pricing Manager - AI Security

Full Time United States $155K - $272K per year Software Development

Senior Backend Engineer (Java / Kotlin / AWS)

Full Time Poland Software Development

Clinical Quality Assurance Coordinator (32949)

Full Time United States $30 - $33 per hour Software Development

Founding Account Executive, SaaS (UK & Europe)

Full Time United Kingdom Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 219,827+ Jobs in Data Engineer

Answer easy questions

Answer easy questions

219,827+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified