You will own the semantic data layer and retrieval pipelines to ensure the accuracy and reliability of the AI platform. You will also act as the primary technical liaison between product engineering and the data team to translate business requirements into actionable data structures.
About the role
Stream Companies is a full-service marketing agency building software for retail automotive dealerships. We operate an AI assistant built on Snowflake Cortex as part of our OrangeOS product suite.
When our AI assistant returns a wrong answer, the issue is almost always upstream—a frozen dealer feed or a mismatched metric definition. As Lead Analytics Engineer, you will take full ownership of the semantic data layer beneath the AI, ensuring complete data accuracy, pipeline reliability, and platform performance.
What you'll do
Semantic Layer & Snowflake Governance
Administer Snowflake roles, grants, and objects supporting the AI platform (warehouse-level decisions remain with the central data team)
Gain fluency in our warehouse and navigate it effectively across platform work
Own the semantic views that encode business meaning: what counts as a sale, what counts as a lead, the grain each metric lives at, etc. An error there produces the same error in every downstream answer
Audit, refactor, and rebuild existing data models during your first 60–90 days
Retrieval quality & pipeline reliability
Own the levers that determine retrieval accuracy: chunk size, document parsing, metadata design, and embedding configuration
Test chunking approaches against dealer paperwork rather than assuming defaults. Snowflake recommends chunks under 512 tokens as a starting point
Own the ingestion pipelines feeding the platform: inventory, dealer & OEM feeds, CRM & DMS data, and marketing performance
Monitor freshness and catch failures, so a broken feed reaches you as an alert before it reaches a client as a wrong number in the platform
Accuracy and compliance
Reconcile AI outputs against core reporting models to guarantee data precision
Audit vehicle offers, payment calculations, and incentives to mitigate advertising compliance risk
Expand evaluation test sets, turning output failures into specific data corrections
Track inference and warehouse cost-per-dealer to ensure feature scalability
Product engineering and the data team
Work with platform engineers on how semantic layer output is queried, cached, and surfaced: what the API returns, how a metric renders in the interface, and what the product shows when a value is null or a feed is stale
Review schema changes with backend engineers before they ship, and work through with front-end engineers how a number should be labeled and qualified on screen
Act as the standing point of contact between product engineering and the data team, carrying platform requirements that need warehouse-level work to them and translating their constraints into decisions product can act on
Route data defects surfaced through the assistant to the data team with enough detail to be actionable, rather than a ticket reporting that the AI was wrong
Partner delivery, stakeholders, and cost
Act as our technical counterpart on partner-built work: define the acceptance criteria, review deliverables against them, and operate the system unassisted before an engagement closes
Ensure all work is easily administered and explainable
Convert business questions from product and account teams into data structures within OrangeOS
Track inference and warehouse cost per dealer, which determines whether a feature is viable at scale
Qualifications
Required
Snowflake Administration: Direct experience managing RBAC, roles, and grants (not query-only access)
Semantic Modeling: 3+ years building metric layers on cloud data warehouses (Snowflake semantic views, dbt metrics, LookML, Cube, etc.) with deep SQL expertise
Pipeline Ownership: Proven experience maintaining production pipelines, monitoring failures, and inheriting legacy models
Python Skills: Fluent in Python for data manipulation, API integrations, and automation
Applied GenAI Experience: Hands-on experience with LLM APIs, prompt engineering, tool call integration, and basic retrieval pipelines (independent projects count)
Communication: Ability to clearly explain AI data behavior to non-technical executives
Preferred
Snowflake Cortex specifically: Analyst, Search, or Agents
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”