Match your resume skills with our AI powered skill match!
You will build and maintain scalable data pipelines that integrate LLMs for extraction, enrichment, and validation of billions of records. You are responsible for ensuring data reliability, cost management, and observability across the entire data platform.
Full-time · Remote (India or Vietnam)
Team: Data Platform
Reporting: Data Platform Lead
Firmable is the market-leading B2B sales intelligence platform in Asia Pacific — and we're scaling that success globally at pace. Backed by leading investors and 2,000+ customers strong, we exist to give sales teams an unfair advantage: the deepest company and people data of any platform, enriched with real-time signals, served at the right moment by intelligent agents.
Our data is the product. Building it now means building with LLMs in the pipeline, and knowing exactly when to trust them.
As Data Engineer — Data Platform, you'll build the pipelines that turn billions of raw records into the clean, modelled datasets our product, analytics, and AI systems depend on. Some of that is classic data engineering: SQL, dbt, orchestration, warehouse performance. A growing share is not: LLM-based extraction, enrichment, and validation running inside the pipeline at scale.
The hard part isn't calling a model. It's making a probabilistic component behave like a reliable one: labelled eval sets, precision and recall you can measure, versioned prompts, cost ceilings, and the harnesses that catch regressions when a vendor silently changes the model under you.
This is a hands-on engineering role with real ownership. You'll own reliability, cost, and whether downstream consumers can trust the data, whether it came from a SQL join or a language model.
LLM-in-the-pipeline systems — extraction, enrichment, entity resolution, and semantic validation steps that run over millions of records a day with structured outputs, retries, and human-review fallbacks
Eval harnesses — labelled sets, scorers, and regression suites that gate every prompt or model change; precision/recall tracked per check, not vibes
Rules vs. LLMs — deterministic checks (dbt tests, data contracts, SQL) wherever structure allows; LLMs where semantic judgement is needed; the discipline to know which is which
Cost and drift — token budgets per pipeline, model routing (cheap models for classification, stronger ones for hard cases), drift detection on vendor updates
Observability — every LLM call logged with prompt version, model, cost, latency, and decision, alongside standard pipeline alerting that surfaces problems before they cascade
Core pipelines and warehouse — Airflow orchestration, dbt models across staging to mart, Snowflake performance and cost over billions of rows, AWS infrastructure as code
Matching and deduplication — embeddings and retrieval patterns for company and people entity resolution across 13 markets
Must Haves
3+ years in data engineering or data infrastructure, with production pipelines you've owned end to end
Strong Python and SQL — production-grade, performance-aware, comfortable at very large scale
Shipped LLMs inside data pipelines — extraction, enrichment, or validation in production, with structured outputs and a labelled eval set that tells you where the model gets it wrong
Harness engineering experience — you've built or maintained eval or test harnesses for LLM outputs, and you can talk through what they caught
Sharp judgement on rules vs. LLMs — you reach for a regex or a dbt test first and can defend the call either way
Production dbt and Airflow — modular, tested models; DAGs that recover gracefully
Cloud warehouse experience — Snowflake preferred; schema design, query optimisation, cost management
AI coding tools are how you work — Claude Code, Cursor, or equivalent, daily, with real shipped work to show for it
Ownership and systems thinking — you weigh upstream dependencies and downstream impact before changing anything
Highly Valued
Eval frameworks and LLM tracing (Logfire, OpenTelemetry)
Embeddings, vector search, or fuzzy matching for entity resolution at scale
Fine-tuning or distilling small models to replace expensive LLM calls
AWS at scale (S3, Lambda, Glue, ECS, RDS) and PostgreSQL
Spark or PySpark; streaming (Kafka, Kinesis, Snowpipe Streaming)
Data privacy and compliance (GDPR, SOC2, CCPA)
Firmable is an AI-native organisation. AI coding tools, automated testing, and AI-assisted review are how we work by default. Every LLM check ships with a labelled eval set, measured precision/recall, and a prompt version you can roll back. Every LLM call in production is logged from day one; retrofitting later is not the plan.
We run lean and ship fast — small senior teams, no layers, minimal process, weekly releases moving toward daily. Teams own their stack end to end. There are no fixed hours and no handholding. If you're not already working this way, this role isn't right for you.
LLMs as production infrastructure — not a demo, not a notebook; models making millions of decisions a day on data customers pay for
Greenfield harnesses — eval coverage, drift detection, and cost controls for in-pipeline LLMs are largely unbuilt; you'll ship them
Scale that matters — billions of rows, 13 markets, and a dataset nobody else has
Small team, massive leverage — your pipelines reach every Firmable customer, every day
Competitive base + meaningful equity — a share in the upside we're building toward
Firmable is an equal opportunity employer. We believe diverse teams build better products.
Ready to build the AI-native data platform behind the world's smartest B2B sales intelligence platform? Apply now — let's talk!
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Data Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”