Design and optimize production-grade RAG pipelines, including document ingestion, semantic enrichment, chunking, hybrid retrieval, reranking, and context-window management. Build semantic caching and monitor retrieval accuracy, hallucination rates, latency, and token costs to improve grounded AI responses.
This is a remote position.
RAG Architect
Job Details
Employment Type: Contract
Work Mode: Remote
Location: Offshore
Total Experience Required: 6 to 10 years
Relevant Experience Required: 3+ years of dedicated experience designing production-grade Retrieval-Augmented Generation (RAG) architectures and optimizing LLM token throughput
Mandatory Certification: Google Cloud Certified Professional Cloud Database Engineer, AWS Certified Data Analytics - Specialty, or Databricks Certified Data Engineer Professional
Job Summary
We are seeking an experienced Context Window Optimization / RAG Architect to take full ownership of our enterprise generative AI retrieval performance, accuracy, and operational cost metrics. The ideal candidate will design high-throughput knowledge retrieval systems, optimize semantic context parsing, build custom re-ranking pipelines, and engineer caching grids to deliver data-grounded AI responses with minimal latency and maximum token efficiency.
Key Responsibilities
Architect end-to-end advanced Retrieval-Augmented Generation (RAG) pipelines, building structures for document parsing, semantic metadata enrichment, and multi-vector lookups.
Optimize context window utilization patterns, designing smart parent-child chunking models, sentence-window retrievals, and sliding window strategies to eliminate irrelevant text tokens.
Build high-performance re-ranking layers, deploying machine learning cross-encoders (e.g., Cohere Rerank, BGE-Reranker) to score retrieved documents before feeding them into the LLM context pool.
Implement automated semantic caching architectures, utilizing caching layers (e.g., GPTCache) to capture recurring semantic queries, reducing API token expenditures and response latencies.
Establish automated data chunking pipelines, configuring ingestion routines to cleanly parse semi-structured and unstructured formats (PDFs, corporate wikis, SQL outputs) into clean vector targets.
Govern vector similarity spaces, fine-tuning hybrid search algorithms that cleanly combine dense semantic embeddings with sparse keyword token indexes (BM25).
Audit context-level hallucination rates and accuracy logs, tracking precision metrics, retrieval recall bounds, and processing speeds to systematically eliminate incorrect model generations.
Requirements
6 to 10 years of enterprise data engineering, database design, or search engine engineering experience, with 3+ dedicated years actively scaling context retrieval loops for live LLM applications.
Strong technical mastery of Python, vector databases (Pinecone, Milvus, Weaviate), text embedding models, open-source orchestration tools (LlamaIndex, LangChain), and SQL.
Deep structural understanding of context window limitations ("lost in the middle" phenomenon), multi-modal token dynamics, network data transfer speeds, and cloud memory spaces.
Mandatory certification: Professional Cloud Data/Database Engineer or Specialty Analytics credential from a major cloud vendor (AWS/GCP/Azure).
Preferred Qualifications
Prior experience implementing Graph RAG frameworks utilizing native knowledge graphs (e.g., Neo4j) to map complex corporate data relationship networks.
Familiarity with fine-tuning open-source text embedding models specifically optimized for industry-specific terminology or legacy product schemas.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”