For Employers

Firmable

Lead Data Engineer - Sourcing

Posted 2 days ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review

Questions interviewers often ask for this role, with sample answers.

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

Architect and own the end-to-end extraction and ETL pipeline to transform unstructured web data into a high-accuracy B2B dataset. Define architectural standards for extractor frameworks, LLM infrastructure, and agentic workflows while ensuring cost-effective and reliable data processing.

Firmable is the market-leading B2B sales intelligence platform in Asia-Pacific. Our competitive moat is our data: the deepest, most localised company and people dataset in every market we operate in. We've proven the model in ANZ. Now we're scaling it across Southeast Asia and the US.

Every record Firmable sells starts in a sourcing pipeline. This role architects that pipeline.

The Role

Lead Data Engineer: Sourcing designs and owns the end-to-end extraction and ETL pipeline that turns unstructured web data into the world's most accurate B2B dataset across 13 markets. You set the architectural standards: extractor patterns, proxy strategy, LLM infrastructure as production systems, agentic escalation workflows, and cost-aware orchestration.

This is not a pipeline maintenance or incremental optimisation role. You're architecting the framework other engineers ship into, building for the hardest extraction problems (anti-bot defences, JS-heavy rendering, schema drift), and owning the sourcing layer end to end.

~80% hands-on architecture and reference implementation: designing extractor frameworks, shipping the hardest extractors, building agentic pipelines, writing eval harnesses and labelled datasets. ~20% cross-functional voice: technical input on sourcing decisions with data platform, product, and analytics; setting standards for coverage, schema, and accuracy.

What You'll Own

Sourcing and ETL architecture

  • End-to-end pipeline design: extraction, normalisation, deduplication, validation, and load; plus cost, performance, and reliability of the whole layer

  • Sourcing layer standards: coverage, accuracy, schema design; the rule-vs-LLM standard codified so the team applies it without you

  • Cost and performance optimisation: cost ceilings with the token math behind them, incremental processing, recovery design, scheduling

Extraction systems and frameworks

  • Extractor framework: patterns, abstractions, and tooling other engineers ship into; new extractors are fast to build and reliable to run

  • Hard-source extractors: anti-bot defences, JS-heavy rendering, schema drift, low-quality structure; plus the proxy and IP rotation strategy behind them

  • Agentic extraction pipelines: rule-based triage, LLM escalation, structured-output validation, retries, human-review queues

LLM infrastructure and observability

  • LLMs as production systems: versioned prompts, labelled eval sets, measured precision and recall, prompt versioning you can roll back, judges debugged on real data

  • Eval and observability scaffolding: eval frameworks, prompt versioning, traces, drift detection when a vendor silently updates a model

  • Model-choice playbook: cheap models for classification, stronger models for nuanced extraction, frontier models for hard edge cases; revised as model economics shift

Skills library and orchestration

  • Skills library: versioned SKILL.md specs for recurring extraction patterns, invocable by engineers and agents alike

  • Orchestration: Airflow or equivalent patterns that scale with data volumes and source counts; dependency management, recovery, cost-aware scheduling

What We're Looking For

Must Haves

  • 6+ years shipping production extraction, ETL, or data pipeline systems in business-critical environments; deep web extraction at scale with experience designing around anti-bot defences, proxy architecture, JS rendering, schema drift, and recovery

  • Shipped LLMs inside extraction pipelines as production systems: structured outputs, versioned prompts, labelled eval sets, logged traces, judges debugged on precision/recall, drift detected on vendor model updates. You can show the repo.

  • Built agents and tool-calling pipelines: architected agent workflows, written SKILL.md specs others depend on, run tool-calling systems in production

  • Sharp judgement on rules vs. LLMs: reach for deterministic logic first, defend the call either way, and have codified the standard for a team

  • Expert Python: production-grade, performance-aware, comfortable with concurrency and scale; plus advanced SQL for complex transformations and performance optimisation

  • Extensive Airflow or equivalent: orchestration, dependency management, recovery patterns, cost optimisation at production scale

  • Shipped with agentic IDEs: Claude Code, Cursor, or equivalent as your default mode, with shipped extraction systems to show for it

  • Architecture judgement and a product mindset: when to refactor, optimise, ship, or start again; you care about coverage and accuracy landing with customers, not uptime metrics

  • You live and breathe AI tools. Structured outputs, evals, traces, and LLM tracing (Logfire, OpenTelemetry, similar) as your default way of working, not a productivity experiment. In 2026, this is how data engineers build extraction at scale and you need to already be doing it.

Highly Valued

  • Cloud data platforms: Snowflake, Redshift; AWS for pipeline deployment (Lambda, S3, ECS, Glue)

  • Vector databases, embeddings, or retrieval patterns for matching and deduplication

  • Eval frameworks like Braintrust, Promptfoo, or Inspect; LLM tracing tooling

  • Data quality frameworks with automated testing and anomaly detection at scale

  • B2B data: firmographics, people data, company registries across markets

  • Data privacy and compliance: GDPR, CCPA

  • Startup or scaleup experience where you shipped fast and owned outcomes end to end

The Environment

Firmable runs lean and ships fast: small senior teams, no layers, minimal process. Sourcing is a core competitive advantage; this role sets the pace for how fast and accurate the entire data pipeline moves. Weekly releases moving toward daily; no fixed hours, full autonomy on architecture choices.

We are an AI-native organisation. That means AI isn't a tool we reach for: it's the default operating mode. Agentic extraction pipelines, LLM-powered classification with eval sets, structured-output validation, labelled datasets versioned with prompt runs, traces logged from day one with cost and latency baked in. If you're not already working this way, this role isn't right for you.

Why This Role

  • Own the sourcing engine: every record Firmable sells starts in the pipeline you architect; your design shapes the accuracy and cost of every customer interaction

  • Greenfield AI-native architecture: the extractor framework, skills library, eval harnesses, agentic orchestration, and production LLM infrastructure are largely unbuilt; you'll architect them from first principles

  • Frontier technical problems: agentic extraction at production scale, LLM-as-extractor with full eval coverage, vendor model drift detection, rule-vs-LLM orchestration across 13 markets

  • Small team, massive leverage: your architecture reaches every Firmable customer, every day; you write the reference implementation daily and own outcomes end to end

  • Real career runway: this role is a bottleneck role; progress is measured by pipeline speed and accuracy, not tenure

Who This Role is NOT For

  • Anyone looking to optimise an existing pipeline rather than architect from first principles

  • Anyone who treats infrastructure as someone else's problem; sourcing architecture lives in your code every day

  • Anyone who sees LLMs as a parsing shortcut rather than production systems that need eval sets, versioning, and drift detection

  • Anyone uncomfortable shipping with minimal process or without an established playbook to follow

  • Anyone who treats AI as something they'll learn on the job rather than already use daily in infrastructure work

  • Anyone coming from a pure ops or data analytics background without shipped systems engineering at scale

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Network Cloud Security Specialist

Full Time Germany Software Development

Remote Satellite Systems I&T Engineer

Freelance United States $50 - $70 per hour Software Development

Solution Architect AI & ML for Life Sciences

Full Time United States $119K - $221K per year Software Development

Salesforce QA Analyst - LATAM

Full Time, Freelance Brazil Software Development

Principal Marketing Analytics Consultant

Full Time United States $70000 - $110K per year Software Development

Full Stack Developer

Full Time United States $110K - $114K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 217,169+ Jobs in Data Engineer

Answer easy questions

Answer easy questions

217,169+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified