Build and operate automated machine learning and generative AI pipelines for training, deployment, serving, and monitoring models in secure, scalable container environments. Manage model and data versioning, performance and drift monitoring, inference optimization, LLMOps infrastructure, and access controls.
This is a remote position.
AI / ML Ops Engineer
Job Details
Employment Type: Contract
Work Mode: Remote
Location: Offshore
Total Experience Required: 4 to 8 years
Relevant Experience Required: 3+ years of dedicated MLOps or DevOps experience deploying and managing machine learning models in production
Mandatory Certification: AWS Certified Machine Learning - Specialty, Google Cloud Certified Professional Machine Learning Engineer, or Databricks Certified Machine Learning Professional
Job Summary
We are seeking an experienced AI / ML Ops Engineer to bridge the gap between data science research and cloud infrastructure execution. The ideal candidate will build automated pipelines to train, test, deploy, and monitor machine learning models and generative AI systems securely and at scale within enterprise container environments.
Key Responsibilities
Design and automate end-to-end ML pipelines (Continuous Training and Continuous Deployment) using orchestration engines like Kubeflow, MLflow, or AWS SageMaker Pipelines.
Orchestrate containerized model deployments, configuring low-latency inference endpoints, auto-scaling GPU/CPU clusters, and model serving runtimes on Kubernetes (KServe, Triton Inference Server).
Implement robust model tracking and data versioning foundations, managing feature stores (e.g., Feast), model registries, and version controls for massive datasets using DVC.
Build automated AI performance and data monitoring gates, tracking model accuracy decay, data drift indicators, concept drift parameters, and system processing latencies in real time.
Optimize inference execution environments, leveraging model compilation engines (e.g., ONNX, TensorRT) and quantization strategies to shorten response times and minimize cloud compute costs.
Integrate generative AI and LLM operational frameworks (LLMOps), configuring semantic caching layers, vector database scaling parameters (e.g., Pinecone, Milvus), and prompt validation pipelines.
Govern machine learning access controls and security profiles, configuring strict data separation barriers, model access tokens, and encryption protocols to safeguard sensitive inference logs.
Requirements
4 to 8 years of core software engineering, DevOps, or data engineering experience, with 3+ dedicated years actively building and maintaining MLOps automation infrastructures.
Strong technical mastery of Python programming, container orchestration (Docker, Kubernetes), ML frameworks (PyTorch, TensorFlow, Hugging Face), and advanced SQL.
Deep structural understanding of distributed system mechanics, GPU resource management limits, model deployment patterns (Shadow, Canary, A/B), and cloud provider API governance.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”