Match your resume skills with our AI powered skill match!
You will design and operate scalable cloud infrastructure, focusing on ML and AI workloads including inference serving and automated deployment pipelines. You are responsible for maintaining platform reliability, cost efficiency, and implementing robust observability and GitOps practices.
We're looking for a Platform Architect who can set the standard for how we build, ship, and operate reliable cloud platforms at scale. You sit at the intersection of platform engineering and SRE. You'll own the path from infrastructure design to reliable production services, bringing DevOps rigor to complex systems.
This is not a ticket-processing role, and it's not a research role. You'll tackle hard problems: platform reliability, scalability, cost efficiency, deployment automation, and workload operations. You'll have the scope to solve them properly. Senior professionals here identify problems before they're asked and raise the ceiling on what the platform can do.
What you will work on
• Build and operate model and inference serving infrastructure, managing latency, throughput, autoscaling, and reliability for real-time and batch inference across multiple tenants.
• Own the ML deployment lifecycle: model registry, versioning, promotion workflows, rollout strategies (canary, shadow, A/B), and safe rollback.
• Operate agentic and LLM workloads in production, managing inference providers and gateways, quota and throttling behavior (TPS/TUPS limits), guardrails, prompt/version management, and graceful degradation under load.
• Build reproducible, automated ML pipelines: training, evaluation, and deployment pipelines as code, with lineage and reproducibility built in.
• Extend infrastructure-as-code to ML systems, using Terraform patterns and multi-project design that bring ML infrastructure under the same standards as the rest of the platform.
• Operate GitOps for ML workloads, owning ArgoCD configuration and promotion workflows across environments and tenants.
• Run ML and AI workloads on multi-tenant Kubernetes (GKE), managing GPU/accelerator scheduling, workload placement, tenant isolation, and cost-aware capacity.
• Own ML reliability and observability: SLOs for inference services, model and data drift detection, performance regression monitoring, alert quality, on-call ergonomics, and runbook culture.
• Drive ML cost efficiency by right-sizing accelerators, managing committed-use and Spot VM capacity, and attributing inference cost across tenants and workloads.
• Use agentic coding tools for infrastructure and pipeline work: scaffolding environments, generating and reviewing IaC and pipeline code, and accelerating automation.
What you won’t find here
A platform team that maintains the status quo. We're actively building: new scale requirements, new architectural domains, and an ML/AI footprint that's growing fast. Senior engineers here shape how the platform evolves, and the tools available to do it are better than they've ever been.
Must have
• 5+ years in platform engineering, SRE, or infrastructure, with meaningful time operating production systems at scale.
• Strong SRE/DevOps foundation. You've owned reliability for production services, defined and measured SLOs, run post-mortems, and driven measurable improvements.
• Deep Terraform expertise. You actively manage complex Terraform state, reusable modules, and multi-project configurations in production, with CI-driven plan/apply workflows.
• Strong GitOps background (ArgoCD or Flux in production). You understand declarative infrastructure management at depth and have opinions on how to do it well.
• Deep Kubernetes knowledge. You've operated clusters in production, dealt with real failure modes, and understand the system at the control plane level.
• Strong cloud infrastructure background across at least one major public cloud (AWS, Azure, or GCP), including networking, compute, IAM, storage, and multi-account or multi-project design.
• Hands-on experience building and operating CI/CD pipelines (GitHub Actions, Cloud Build, GitLab CI, or equivalent).
• Automation-first thinking at a senior level. You implement systems that eliminate entire categories of manual work.
• Active user of agentic coding tools. You know how to direct them effectively, review their output critically, and use them to multiply your output.
• Strong communicator. You can articulate operational decisions, technical trade-offs, and incident summaries clearly to engineers and leadership alike.
Nice to have
• MLOps experience, including hands-on experience deploying and operating ML or AI workloads in production: serving, inference, or training infrastructure that real users depended on.
• Strong GCP experience, including VPC networking, Compute Engine, IAM, Cloud Storage, multi-project or organization design, and GKE (Standard and/or Autopilot).
• Hands-on experience with GCP data services, especially BigQuery in production: partitioning and clustering, query cost and performance tuning, and dataset-level IAM. Familiarity with at least one of Dataflow, Pub/Sub, or Dataproc.
• Experience with GPU/accelerator scheduling and node lifecycle management in production (e.g., GKE node auto-provisioning, GPU time-sharing, or equivalent).
• Experience operating LLM inference at scale, managing provider quotas/throttling (TPS/TUPS), gateways, caching, and guardrails (e.g., Vertex AI, Gemini API, or equivalent).
• Experience with ML pipeline and orchestration tooling such as Argo Workflows, Kubeflow, Cloud Composer/Airflow, Vertex AI Pipelines, or equivalent.
• Experience with model registries, feature stores, and experiment tracking (e.g., MLflow, Feast, or equivalent).
• Familiarity with model and data drift monitoring and ML-specific observability.
• Background in FinOps: inference cost attribution, committed use discount (CUD) and reservation planning, and accelerator capacity forecasting.
• Familiarity with data infrastructure such as object storage, CDC pipelines, or lakehouse patterns.
• Experience with multi-tenant infrastructure: isolation patterns, noisy neighbor mitigation, and tenant lifecycle management.
• Prior experience scaling ML or platform infrastructure at a startup moving toward enterprise-grade requirements.
Location: Remote in LATAM
Payment in USD
Working hours: EST time zone
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Architect
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”