AI Engineer (Data Guardrails & LLM Ingestion Pipelines)

 Posted 11 hours ago
  
 Worldwide
  
 $15000 - $20000 per year
  
2-5 years experience
Apply Now

Please mention DailyRemote when applying

AI Summary

Design and build scalable data ingestion, cleaning, and LLM enhancement pipelines to power AI applications. Implement AI guardrails and RAG architectures to ensure the accuracy and reliability of AI-ready datasets.

This is a remote position.

About the Role

We are looking for a highly skilled AI Engineer to design and build robust data ingestion, cleaning, validation, and LLM enhancement pipelines that power our AI applications. You will transform raw, unstructured data into high-quality, AI-ready datasets while implementing guardrails that ensure accuracy, consistency, and reliability.

Key Responsibilities

·         Design and develop scalable data ingestion pipelines for structured and unstructured data.

·         Build automated data cleaning, normalization, and preprocessing workflows.

·         Develop AI-powered enrichment pipelines using LLMs (OpenAI, Claude, Gemini, etc.).

·         Implement data quality validation and AI guardrails.

·         Develop prompt engineering workflows for data transformation.

·         Build document processing pipelines for PDFs, Word documents, CSVs, websites, and APIs.

·         Develop Retrieval-Augmented Generation (RAG) pipelines.

·         Create evaluation frameworks for LLM quality and accuracy.

·         Build ETL/ELT workflows for AI-ready datasets.

·         Integrate vector databases for semantic search.

·         Monitor pipeline performance, cost, latency, and data quality.

·         Collaborate with cross-functional teams to deliver production AI systems.

Required Technical Skills

Programming

·         Python (Expert)

·         SQL

·         Git

AI & LLMs

·         OpenAI API

·         Anthropic Claude API

·         Google Gemini API

·         Prompt Engineering

·         Function Calling

·         Structured Outputs

AI Frameworks

·         LangChain

·         LlamaIndex

·         DSPy (Preferred)

·         PydanticAI (Nice to Have)

Data Engineering

·         Pandas

·         Polars

·         ETL/ELT Pipelines

·         Apache Airflow (Preferred)

·         Data Validation Frameworks

Vector Databases

·         Pinecone

·         Weaviate

·         Qdrant

·         ChromaDB

·         FAISS

Cloud & Infrastructure

·         Docker

·         Kubernetes (Preferred)

·         AWS / Azure / GCP

·         Linux

Databases

·         PostgreSQL

·         MongoDB

·         Redis


Requirements

Preferred Qualifications

·         Experience building production-grade AI systems.

·         Strong understanding of RAG architectures.

·         Experience implementing AI guardrails and hallucination mitigation.

·         Experience with OCR and document parsing.

·         Experience with embedding models and semantic search.

·         Knowledge of data governance and security best practices.

Success Metrics

·         Build scalable ingestion pipelines.

·         Deliver automated data cleaning and LLM enhancement workflows.

·         Implement AI guardrails to improve output quality.

·         Develop evaluation pipelines for LLM performance.

·         Contribute to a production-ready AI platform.


Similar Jobs

See all Remote Software Development jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in AI Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified