Design, build, and stabilize production-grade multi-agent systems that coordinate specialized agents, manage state and shared memory, and execute enterprise API, database, and file operations securely. Implement human approval checkpoints, advanced reasoning and error-correction patterns, and observability tools to monitor agent behavior and diagnose failures.
This is a remote position.
Agentic Workflow / Multi-Agent Systems Engineer
Job Details
Employment Type: Contract
Work Mode: Remote
Location: Offshore
Total Experience Required: 4 to 8 years
Relevant Experience Required: 2+ years of dedicated experience developing multi-agent systems, autonomous AI agents, and complex task-planning state machines
We are seeking an experienced Agentic Workflow / Multi-Agent Systems Engineer to design, develop, and stabilize autonomous multi-agent networks within our enterprise environment. The ideal candidate will move past single-prompt solutions to build production-grade, stateful multi-agent systems where distinct specialized AI entities collaborate, share context, call enterprise APIs, handle execution failures gracefully, and execute multi-step business operations independently.
Key Responsibilities
Design and develop multi-agent orchestration architectures using frameworks like LangGraph, CrewAI, AutoGen, or Semantic Kernel.
Build stateful deterministic and non-deterministic state machines, managing conversation loops, task delegation logic, branching paths, and agent-to-agent communication networks.
Implement secure, robust tool execution frameworks, empowering agents to dynamically call enterprise REST APIs, query databases via SQL, and process files within sandboxed runtime execution environments.
Configure sophisticated multi-agent memory layers, setting up short-term transactional memory, long-term semantic vector storage, and cross-agent context sharing protocols.
Establish structured human-in-the-loop (HITL) validation checkpoints, configuring human approval gates for sensitive agentic operations like financial triggers or external data mutations.
Optimize agent reasoning and planning structures, implementing advanced cognitive patterns such as Reason and Act (ReAct), Plan-and-Solve, and self-reflection error-correction loops.
Implement agentic observability and tracing frameworks using platforms like LangSmith, Arize Phoenix, or Datadog LLM Observability to diagnose stuck loops, trace call sequences, and monitor token consumption.
Requirements
4 to 8 years of core enterprise backend web engineering, asynchronous programming, or distributed systems experience, with 2+ dedicated years actively writing production-level application code for autonomous multi-agent environments.
Strong technical mastery of Python or TypeScript, asynchronous programming (Asyncio), state-management patterns, API integration design, and relational databases.
Deep structural understanding of LLM cognitive limits, tool calling hallucination profiles, infinite loop mitigation strategies, token budget planning, and multi-agent consensus protocols.
Mandatory certification: Professional ML Engineer or Specialty Machine Learning credential from a major cloud vendor (AWS/GCP/Azure).
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”