Design and deploy production-grade LLM applications, RAG pipelines, and multi-agent workflows that connect enterprise data and systems with foundation models. Build secure, observable, and cost-efficient middleware and integrations, including safeguards, evaluation, caching, and performance monitoring.
This is a remote position.
LLM Integration / LangChain Engineer
Job Details
Employment Type: Contract
Work Mode: Remote
Location: Offshore
Total Experience Required: 4 to 8 years
Relevant Experience Required: 2+ years of dedicated hands-on experience building, integration, and deploying applications powered by Large Language Models (LLMs)
We are seeking an experienced LLM Integration / LangChain Engineer to design, develop, and implement the orchestration layers connecting our enterprise data assets with cutting-edge generative AI models. The ideal candidate will build production-grade Retrieval-Augmented Generation (RAG) pipelines, program multi-agent reasoning loops using LangChain, LangGraph, or LlamaIndex, and establish secure system middleware to safely deploy AI capabilities at scale.
Key Responsibilities
Design and construct advanced LLM applications and orchestrations using specialized frameworks like LangChain, LangGraph, LlamaIndex, or AutoGen.
Develop complex multi-agent reasoning chains and workflows, implementing custom tool calling structures, memory caching architectures, and guardrail validations.
Expose and consume programmatic endpoints, constructing high-throughput API integrations connecting foundational LLMs (e.g., OpenAI, Anthropic, open-source models via Hugging Face/Ollama) with internal corporate databases and CRMs.
Apply rigorous AI evaluation and prompt tracking structures, utilizing observability platforms (e.g., LangSmith, Arize Phoenix) to monitor token usage bounds, model latency, and prompt generation drift.
Implement secure middleware execution barriers, configuring text sanitization, PII data-masking pipelines, prompt injection defensive rings, and toxicity filtering parameters.
Optimize model inference costs and context window budgets, designing custom semantic caching frameworks (e.g., GPTCache) to intercept recurring operational queries.
Requirements
4 to 8 years of core enterprise backend web engineering or data pipelines experience, with 2+ dedicated years actively writing production-level application code wrapped directly around LLM infrastructures.
Strong technical mastery of Python or TypeScript, vector representations, prompt engineering grounding mechanics, asynchronous web frameworks (FastAPI), and SQL.
Deep structural understanding of transformer model designs, text embedding properties, agentic tool execution cycles, and API orchestration limits.
Mandatory certification: Professional-level machine learning or cloud developer certification from a major cloud vendor (AWS/GCP/Azure).
Preferred Qualifications
Prior experience fine-tuning open-source LLMs (e.g., Llama, Mistral) via quantization techniques like QLoRA or LoRA frameworks.
Familiarity with deploying AI applications within container systems (Docker, Kubernetes) integrated into modern DevSecOps CI/CD delivery loops.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”