For Employers
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

The Senior Data Engineer will own the end-to-end data substrate, including ingestion, modeling, and governance within a Databricks and Azure environment. They will build and maintain serving layers for AI agents while ensuring data quality, performance, and cost-efficiency through automated infrastructure and monitoring.

About Sciene
At Sciene, the mission is to empower professional services firms with cutting-edge, customized AI solutions — enhancing automation, analytics, and optimization across industries while prioritizing security, cost efficiency, and state-of-the-art technology.
Our flagship product, the Sciene AI Companion, is an autonomous customer success platform deployed across Quartile — the world's largest retail media optimization platform, managing performance marketing for 1,000+ brands. It automates relationship-heavy enterprise workflows end to end: generating personalized email replies in the CSM's own voice (8x faster), building full presentation decks for client meetings (12x faster), and detecting and diagnosing account fluctuations before anyone has to ask (6x faster). None of this replaces human judgment — it removes the work that was getting in the way of it.
Read more about how we built it: Sciene AI Companion: Building an Autonomous Customer Success Platform on Databricks

OVERVIEW:

Every answer the AI Companion gives is only as good as the data underneath it. Sciene runs a production lakehouse on Databricks and Azure that ingests advertising, CRM, and conversational data from a dozen upstream systems, curates it under Unity Catalog governance, and serves it back to autonomous agents as tools — not as dashboards.

The Senior Data Engineer owns that substrate end to end: ingestion from messy third-party APIs, modeling and governance in Unity Catalog, the serving layer agents query at runtime, and the health monitoring that tells us something broke before a CSM finds out in a client meeting. Infrastructure is code, deploys go through CI/CD, and quality is measured — not assumed. You write the Terraform, you own the alerts, you get paged by your own monitors.

This is also a role where you work with AI, not just for it. Our engineers ship with coding agents (Claude Code and our MCP-integrated toolchain) as part of the normal workflow, and your consumer is an autonomous agent making decisions on top of what you produce — which raises the bar on correctness, freshness, and how the data is shaped.


REQUIREMENTS:

  • Strong software engineering fundamentals in Python and advanced SQL — this role builds tested, deployed, version-controlled pipelines, not one-off notebooks
  • Deep hands-on experience with Databricks: Spark/PySpark, Delta Lake, Unity Catalog, Workflows, and SQL warehouses, including performance tuning of real workloads
  • Proven experience designing data pipelines and data models in production: medallion or equivalent layered architectures, incremental processing, CDC, and dimensional modeling
  • Experience integrating third-party APIs at scale, with the operational maturity that implies (auth, pagination, throttling, partial failure, backfills, idempotency, schema evolution)
  • Familiarity with cloud platforms (we run on Azure — ADLS, Container Apps, Key Vault, API Management, Entra ID)
  • Experience with Infrastructure as Code (Terraform) and CI/CD pipelines, plus solid git and code-review habits
  • A working definition of data quality: tests, expectations, monitoring, and the discipline to catch problems upstream rather than explaining them downstream
  • Practical experience using AI coding assistants/agents in real engineering work, and the critical eye to know when their output is wrong
  • Excellent problem-solving and analytical skills, and the autonomy expected of a senior engineer: you own a problem end to end — from framing to shipped, monitored outcome — and are accountable for the result, not just the merge
  • Bachelor's degree or higher in Computer Science, Information Technology, or a related field — or equivalent practical experience


PREFERRED QUALIFICATIONS:

  • Experience serving data to LLM-powered applications, including retrieval and context design
  • Experience with operational stores alongside the lakehouse (Databricks Lakebase, PostgreSQL, MongoDB)
  • Observability stacks: OpenTelemetry, Grafana, Loki, structured logging
  • Delta Sharing for secure data exchange with clients and partners
  • Domain background in retail media, digital advertising, or marketplace data
  • Databricks or Azure certifications


WHAT YOU’LL DO:

  • Design, build, and operate ingestion pipelines from advertising and business platforms (Amazon Ads, Google Ads, Walmart Connect, Criteo, Salesforce, and internal services) via REST APIs, Delta Sharing, object storage, and event streams
  • Model curated data layers in Delta Lake under Unity Catalog: medallion architecture, incremental and CDC patterns, clear contracts between layers, lineage, and row/column-level governance
  • Build the serving layer for AI: SQL views, Unity Catalog functions, and low-latency stores exposed to agents as tools through the Model Context Protocol (MCP) behind Azure API Management — designing for what an agent needs to reason well, not just for what a BI tool needs to render
  • Own data quality and health monitoring as code: freshness, volume, schema-drift and business-rule expectations, with alerting wired into Grafana and a clear triage path when something fires
  • Ship everything through Terraform and CI/CD — Databricks and Azure resources, jobs, permissions, and environment promotion from dev to prod, with no manual clicking in the console
  • Own performance and cost: warehouse and cluster sizing, partitioning and liquid clustering, job orchestration in Databricks Workflows / Lakeflow, and continuous scrutiny of what each pipeline actually costs to run
  • Work fluently with AI coding agents as part of your daily engineering practice — writing precise specs, keeping repository context and documentation in a state where agents produce good output, and reviewing what they generate with real judgment
  • Collaborate with cross-functional teams — AI engineers, software engineers, and product — turning product requirements into data the platform can actually serve, and treating internal consumers as customers with contracts and expectations
  • Instrument what you ship with OpenTelemetry and structured logging, and improve reliability and latency in production
  • Document and maintain the codebase, ensuring code quality and adherence to best practices


*This is a PJ contract based in Brazil.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Solutions Engineer

Full Time Thailand Software Development

Senior Full-Stack Engineer

Full Time Puerto Rico Software Development

AI Governance Analyst

Full Time United States Software Development

Senior Data Analyst

Full Time United States Software Development

MEDICAL RECORDS CODER II

Full Time United States Software Development

Project Developer Consultant (1099)

Freelance United States Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Data Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified