This is a remote position.
Senior AI & Agentic System Engineer
Experience: 5+ Years
Location: Gurgaon, India / Fully Remote
Team: AI Platform / Engineering
About the Role
We are looking for a Senior AI & Agentic System Engineer to design and build production-grade AI agents that can reason, use tools, interact with APIs, and execute complex multi-step workflows.
You’ll work at the intersection of LLM engineering, agent architecture, backend systems, and cloud infrastructure, with significant ownership over architecture, reliability, evaluation, and scalability.
Requirements
Key Responsibilities
- Design and build autonomous AI agents capable of tool use, web/API interaction, memory, and multi-step task execution.
- Architect agentic workflows including planner/executor/verifier, multi-agent, HITL, retry, and failure-recovery patterns.
- Integrate LLM providers such as OpenAI, Anthropic, and Gemini, including structured outputs, function calling, context management, and prompt optimization.
- Build RAG and knowledge pipelines using embeddings, vector databases, hybrid search, reranking, and agent memory.
- Develop scalable backend services using Python and/or Node.js/TypeScript, REST/GraphQL APIs, async processing, and background jobs.
- Design reliable distributed systems with appropriate patterns for state persistence, caching, idempotency, event-driven workflows, and fault tolerance.
- Build observability and evaluation frameworks for AI systems, including tracing, behavioral regression testing, latency, quality, and token-cost monitoring.
- Own CI/CD and cloud deployments across development, staging, and production environments.
- Evaluate emerging technologies such as MCP, A2A, reasoning models, and agent frameworks and recommend practical adoption.
Required Skills
- 5+ years of software/backend engineering experience with strong system-design expertise.
- Hands-on experience building production AI agents, beyond basic chatbot implementations.
- Strong experience with at least one major LLM provider SDK: OpenAI, Anthropic, or Gemini.
- Strong understanding of prompt engineering, tool/function calling, structured outputs, context management, and token optimization.
- Experience with agent frameworks/patterns such as LangGraph, CrewAI, AutoGen, ReAct, or plan-and-execute.
- Strong Python and/or Node.js/TypeScript development skills.
- Experience with RAG, embeddings, vector databases, hybrid retrieval, and reranking.
- Strong knowledge of REST/GraphQL APIs, asynchronous processing, databases, and distributed systems.
- Hands-on AWS experience, particularly Lambda, ECS/Fargate, API Gateway, S3, SQS, SNS/EventBridge, IAM, and CloudWatch.
- Experience with Docker, GitHub Actions, and Terraform.
- Understanding of AI security, including prompt injection, guardrails, output validation, and safe tool execution.
AI/LLMOps & Evaluation
Experience with tools such as:
- Langfuse, LangSmith, Helicone, or Arize Phoenix
- RAGAS, DeepEval, Promptfoo, or Braintrust
- LLM tracing, behavioral evaluations, regression testing, and cost monitoring.
Nice to Have
- Browser automation using Playwright or Puppeteer.
- Experience with multimodal/vision-based agents.
- Knowledge of GraphRAG and knowledge graphs.
- Experience with Temporal or durable workflow execution.
- Open-source contributions to AI tooling.
- Background in ML/data science alongside software engineering.
Tech Stack
LLMs: OpenAI, Anthropic Claude
Agent Infrastructure: Custom orchestration, MCP
Backend: Python, Node.js/TypeScript
Cloud: AWS Lambda, ECS/Fargate
Data: PostgreSQL, S3, SQS, DynamoDB
CI/CD: GitHub Actions
Observability: Datadog, CloudWatch, Langfuse
IaC: Terraform
Browser Automation: Playwright