Design and productionize LLM agents and predictive models for an automotive retail platform, focusing on tool-use and real-time customer conversations. Own the full model lifecycle, including evaluation infrastructure, RAG systems, and deployment monitoring.
This is a remote position.
Join one of the Philippines' fastest-growing tech companies! Open to Philippine-based candidates only.
Company Overview
Full Scale is a tech services company that helps businesses build dedicated teams of skilled software engineers. We make finding and retaining experienced software talent easy and affordable.
Position Summary
We are looking for a Senior Data Scientist with an agentic AI focus to join our growing team. You will design, build, and productionize both predictive models and LLM agents for a US-based client in the automotive retail space — a real-time platform where voice calls, customer data, and AI agents come together to power live customer conversations. This is not a notebook-to-engineering-handoff role. The data layer already exists (AWS lakehouse, Databricks, Gold-zone datasets). The ML services already run. We're hiring the scientist who will make them think — building multi-step agents with tool access to dealership systems, owning the evaluation infrastructure, and shipping models that drive real business outcomes.
Key Responsibilities
- Design, build, and evaluate LLM agents that operate against real dealership systems — booking service appointments, answering vehicle availability questions, resolving customer identity, and escalating to humans with full context.
- Own the tool-use layer: define the tools and function schemas agents call, the guardrails around each, and how the system fails when a downstream service is slow or unavailable.
- Build agent evaluation infrastructure — offline eval sets, adversarial and edge-case suites, live A/B testing, and regression gates that block deployment on quality drops.
- Design and implement escalation logic and confidence thresholds: where the agent acts, where it confirms, and where it hands off to a person.
- Develop and maintain RAG systems over dealership content (service history, OEM documentation, policy, inventory) using Bedrock embeddings, pgvector on Aurora, and OpenSearch Serverless.
- Own prompt architecture, versioning, and change control as a first-class engineering artifact under source control.
- Build, validate, and deploy predictive models on lakehouse data — gross profit forecasting, customer lifetime value, defection risk, next-service prediction, identity resolution, and inventory pricing signals.
- Own the full model lifecycle: feature engineering, training, validation, deployment, monitoring, and retraining. Ship models with drift and degradation monitoring from day one.
- Convert business questions from operations and ownership into well-posed modeling problems and push back when a question is better answered with a query than a model.
- Quantify and communicate model impact in dealership terms: gross, units, retention, CSI, labor hours saved.
Requirements
- 4+ years applying data science in production, with models that made real decisions and had real consequences.
- Strong Python and SQL. You write code others can run and maintain.
- Hands-on experience building LLM agents with tool use and function calling — not just prompt engineering. Be ready to walk through a system you built and how you evaluated it.
- Practical RAG experience: embeddings, vector search, chunking, retrieval evaluation, and re-ranking.
- Sound statistical fundamentals and honest handling of uncertainty. We prefer a well-calibrated interval over a confident point estimate.
- Experience deploying models to production — not handing notebooks off to an engineering team.
- Working comfort with AWS ML tooling (Bedrock, SageMaker, Lambda) and a lakehouse or data warehouse environment.
Nice to Have
- Databricks, Spark, and Delta Lake experience.
- Experience with agent frameworks, orchestration patterns, and structured output / JSON-mode reliability.
- Voice AI or conversational systems experience (Retell, Vapi, or comparable), including latency constrained design.
- Time-series forecasting and causal inference exposure.
- Automotive retail domain knowledge — DMS, CRM, F&I, fixed operations. A candidate who understands what an RO or a chargeback is will ramp faster.
- Familiarity with AI safety and evaluation practice, including handling of PII in prompts and logs.
Benefits
Why join us:
- Fully remote – work from anywhere in the Philippines.
- Work on live agentic AI systems — not POCs, not slideware, not handoff-to-engineering.
- Data layer already in place (AWS lakehouse, Databricks) so you can focus on modeling and agents, not plumbing.
- Small, senior, high-autonomy team with documentation-first culture.
- Opportunity to define the evaluation and deployment standards every new model and agent will follow.
- A team environment that values intellectual honesty, technical depth, and follow-through.