For Employers

Firmable

Data Engineering Lead - Data Quality Systems

Posted 2 hours ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

You will lead a team of Applied AI Engineers to architect and maintain systems for measuring, verifying, and enforcing data quality across billions of records. This involves building verification pipelines, LLM-based validators, and release gates to ensure high-quality data reaches customers.

Firmable is the market-leading B2B sales intelligence platform in Asia Pacific — and we're scaling that success globally at pace. Backed by leading investors and 2,000+ customers strong, we exist to give sales teams an unfair advantage: the deepest company and people data of any platform, enriched with real-time signals, served at the right moment by intelligent agents.

Our moat is the data. This role owns whether it can be trusted.

The Role

As Data Engineering Lead — Data Quality Systems, you'll lead a small team of Applied AI Engineers building the systems that measure, verify, and enforce quality across billions of company and people records in 13 markets.

This is not a QA or testing role. You won't be writing ETL test cases, app tests, or product acceptance tests. You'll be architecting the quality layer itself — verification pipelines, eval harnesses, LLM-based validators, anomaly detection, and the release gates that decide what data ships to customers.

~80% hands-on engineering. ~20% leading the team. You set the technical bar, write the hardest code, and unblock the engineers building on top of your frameworks. If you want a pure people-management seat, this isn't it.

What You'll Own

  • Quality systems architecture — design and build the verification, sampling, and scoring pipelines that run continuously over billions of rows across every market

  • Frameworks the team builds on — the harnesses, abstractions, and SKILL.md specs that make new quality checks fast to write, cheap to run, and hard to get wrong

  • Harness engineering — eval harnesses for LLM-based validation and extraction: labelled eval sets, precision/recall tracking, judge calibration, prompt versioning, drift detection on vendor model updates

  • Release gates — pre- and post-production gates that block bad data before it reaches customers, with the failure analysis and triage tooling to match

  • Incident remediation — root-cause quality incidents at scale, ship the fix, and turn the failure into a permanent automated check

  • Team leadership — lead 3–5 Applied AI Engineers: technical direction, code review, pairing, and growing them into engineers who own outcomes end to end

What We're Looking For

Must Haves

  • 7+ years building production data systems in business-critical environments — you've shipped systems that ran unattended, at scale, and stayed up

  • Worked with billions of rows — you know what breaks at that scale, and how to design quality checks that don't

  • Built data quality systems and frameworks — not used them, built them: validation engines, anomaly detection, scoring, sampling strategies, reconciliation against ground truth

  • Harness engineering experience — you've designed eval or test harnesses that other engineers depend on, and you can show the repo

  • Led engineers — you've directed a small team technically, reviewed their code, and stayed hands-on while doing it

  • Strong Python and advanced SQL — production-grade, performance-aware, comfortable with concurrency and large-scale transformations

  • You operate LLMs as production systems — eval sets, versioned prompts, logged traces, cost ceilings, debugged judges on precision/recall

  • Shipped real work with agentic IDEs — Claude Code, Cursor, or equivalent. Not "tried it" — built and shipped with it as your default mode

  • Sharp judgement on rules vs. LLMs — deterministic checks where structure allows, LLMs where semantic judgement is needed, and you can defend the call

Highly Valued

  • B2B data: firmographics, people data, entity resolution, registry matching across markets

  • Cloud data platforms — Snowflake, Databricks, Redshift — and AWS for pipeline deployment

  • Airflow (or equivalent) orchestration at production scale

  • Vector databases, embeddings, or retrieval patterns for matching and deduplication

  • Startup or scaleup experience where you defined the standard rather than inheriting it

How We Build

Firmable is an AI-native organisation. Agentic development, evals, traces, and AI-powered review are the default mode of working, not a productivity experiment. Recurring workflows ship as versioned SKILL.md specs any teammate or agent can run. Every LLM call is logged with prompt version, model, cost, and decision from day one.

We run lean and ship fast — small senior teams, no layers, minimal process, weekly releases moving toward daily. Teams own their stack end to end. There are no fixed hours and no handholding. If you're not already working this way, this role isn't right for you.

Why This Role

  • Own the trust layer — every record Firmable sells passes through systems you architect

  • Greenfield frameworks — the harnesses, gates, and validators are largely unbuilt; you'll define them

  • Frontier problems — LLM-as-judge at production scale, drift detection on vendor models, quality enforcement over billions of rows

  • Small team, real leverage — lead the engineers building the quality standard for 13 markets

  • Competitive base + meaningful equity — a share in the upside we're building toward

Firmable is an equal opportunity employer. We believe diverse teams build better products.

Ready to build the quality engine behind the world's smartest B2B sales intelligence platform? Apply now — let's talk!

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Data Entry jobs →

Administrative Assistant

Full Time United States Data Entry

Administrative Assistant

Full Time United States $77840 - $111K per year Data Entry

Administrative Assistant for an Organizing Business in the US (Home Based Full Time)

Full Time United States Data Entry

Sr Prod Data Input Assoc

Full Time Japan Data Entry

Senior Colleague Technical Consultant (Ellucian Colleague/Data Conversion)

Full Time United States Data Entry

Order Entry Specialist

Freelance United States $16 - $18 per hour Data Entry
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Data Entry

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified