About Sciene
At Sciene, the mission is to empower professional services firms with cutting-edge, customized AI solutions — enhancing automation, analytics, and optimization across industries while prioritizing security, cost efficiency, and state-of-the-art technology.
Our flagship product, the Sciene AI Companion, is an autonomous customer success platform deployed across Quartile — the world's largest retail media optimization platform, managing performance marketing for 1,000+ brands. It automates relationship-heavy enterprise workflows end to end: generating personalized email replies in the CSM's own voice (8x faster), building full presentation decks for client meetings (12x faster), and detecting and diagnosing account fluctuations before anyone has to ask (6x faster). None of this replaces human judgment — it removes the work that was getting in the way of it.
Read more about how we built it: Sciene AI Companion: Building an Autonomous Customer Success Platform on Databricks
OVERVIEW:
Every answer the AI Companion gives is only as good as the data underneath it. Sciene runs a production lakehouse on Databricks and Azure that ingests advertising, CRM, and conversational data from a dozen upstream systems, curates it under Unity Catalog governance, and serves it back to autonomous agents as tools — not as dashboards.
The Senior Data Engineer owns that substrate end to end: ingestion from messy third-party APIs, modeling and governance in Unity Catalog, the serving layer agents query at runtime, and the health monitoring that tells us something broke before a CSM finds out in a client meeting. Infrastructure is code, deploys go through CI/CD, and quality is measured — not assumed. You write the Terraform, you own the alerts, you get paged by your own monitors.
This is also a role where you work with AI, not just for it. Our engineers ship with coding agents (Claude Code and our MCP-integrated toolchain) as part of the normal workflow, and your consumer is an autonomous agent making decisions on top of what you produce — which raises the bar on correctness, freshness, and how the data is shaped.
REQUIREMENTS:
- Strong software engineering fundamentals in Python and advanced SQL — this role builds tested, deployed, version-controlled pipelines, not one-off notebooks
- Deep hands-on experience with Databricks: Spark/PySpark, Delta Lake, Unity Catalog, Workflows, and SQL warehouses, including performance tuning of real workloads
- Proven experience designing data pipelines and data models in production: medallion or equivalent layered architectures, incremental processing, CDC, and dimensional modeling
- Experience integrating third-party APIs at scale, with the operational maturity that implies (auth, pagination, throttling, partial failure, backfills, idempotency, schema evolution)
- Familiarity with cloud platforms (we run on Azure — ADLS, Container Apps, Key Vault, API Management, Entra ID)
- Experience with Infrastructure as Code (Terraform) and CI/CD pipelines, plus solid git and code-review habits
- A working definition of data quality: tests, expectations, monitoring, and the discipline to catch problems upstream rather than explaining them downstream
- Practical experience using AI coding assistants/agents in real engineering work, and the critical eye to know when their output is wrong
- Excellent problem-solving and analytical skills, and the autonomy expected of a senior engineer: you own a problem end to end — from framing to shipped, monitored outcome — and are accountable for the result, not just the merge
- Bachelor's degree or higher in Computer Science, Information Technology, or a related field — or equivalent practical experience
PREFERRED QUALIFICATIONS:
- Experience serving data to LLM-powered applications, including retrieval and context design
- Experience with operational stores alongside the lakehouse (Databricks Lakebase, PostgreSQL, MongoDB)
- Observability stacks: OpenTelemetry, Grafana, Loki, structured logging
- Delta Sharing for secure data exchange with clients and partners
- Domain background in retail media, digital advertising, or marketplace data
- Databricks or Azure certifications
WHAT YOU’LL DO:
- Design, build, and operate ingestion pipelines from advertising and business platforms (Amazon Ads, Google Ads, Walmart Connect, Criteo, Salesforce, and internal services) via REST APIs, Delta Sharing, object storage, and event streams
- Model curated data layers in Delta Lake under Unity Catalog: medallion architecture, incremental and CDC patterns, clear contracts between layers, lineage, and row/column-level governance
- Build the serving layer for AI: SQL views, Unity Catalog functions, and low-latency stores exposed to agents as tools through the Model Context Protocol (MCP) behind Azure API Management — designing for what an agent needs to reason well, not just for what a BI tool needs to render
- Own data quality and health monitoring as code: freshness, volume, schema-drift and business-rule expectations, with alerting wired into Grafana and a clear triage path when something fires
- Ship everything through Terraform and CI/CD — Databricks and Azure resources, jobs, permissions, and environment promotion from dev to prod, with no manual clicking in the console
- Own performance and cost: warehouse and cluster sizing, partitioning and liquid clustering, job orchestration in Databricks Workflows / Lakeflow, and continuous scrutiny of what each pipeline actually costs to run
- Work fluently with AI coding agents as part of your daily engineering practice — writing precise specs, keeping repository context and documentation in a state where agents produce good output, and reviewing what they generate with real judgment
- Collaborate with cross-functional teams — AI engineers, software engineers, and product — turning product requirements into data the platform can actually serve, and treating internal consumers as customers with contracts and expectations
- Instrument what you ship with OpenTelemetry and structured logging, and improve reliability and latency in production
- Document and maintain the codebase, ensuring code quality and adherence to best practices
*This is a PJ contract based in Brazil.