Match your resume skills with our AI powered skill match!
You will own the core infrastructure, production uptime, and scalability of a high-growth AI/ML platform. Responsibilities include managing AWS infrastructure, building CI/CD pipelines, and improving developer productivity through robust tooling and automation.
This is a backend-architecture-heavy Platform Engineer role on a tight-knit engineering team of roughly 15 people, sitting at the intersection of reliability, scalability, and developer experience. You will own the core infrastructure that keeps a fast-moving AI/ML platform running at scale, making it faster, cheaper, and safer for the whole team to build and ship.
Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
Design and improve backend and platform systems for scale, covering capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
Define and continuously improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected and resolved quickly.
Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
Write clean, maintainable code to automate systems, improve backend services, and create internal developer tooling.
2 to 4 years of experience owning production cloud infrastructure for a high-availability, user-facing platform, with direct responsibility for uptime, performance, deployment safety, and cost.
Deep hands-on experience with AWS and containerized systems; strong familiarity with Terraform, Kubernetes/EKS, Docker, EC2, networking, load balancers, and secrets management.
A track record of building or operating CI/CD, release automation, observability, alerting, and incident response systems.
Strong backend engineering judgment across service architecture, APIs, databases, async systems, queues, and production failure modes.
Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services.
Prior exposure to AI/ML platform engineering or software engineering in an AI/ML context; generic DevOps or infrastructure-only backgrounds are not a strong fit.
Experience operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms is a plus.
Demonstrated focus on reducing cloud spend through better architecture, autoscaling, workload placement, caching, or cleanup systems.
High ownership mindset and genuine interest in learning; strong technical aptitude matters more than a specific number of years.
Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
On-site in Singapore. Candidates based in San Francisco are also considered on-site there; fully remote arrangements are available for candidates based in Europe or other regions outside the US and Southeast Asia.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Platform Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”