See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewQuestions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will own the reliability, scale, and performance of core infrastructure systems while building and maintaining AWS-based platform services. Additionally, you will define observability workflows, manage CI/CD pipelines, and write production-quality code to improve developer productivity.
This is a backend-architecture-heavy platform engineering role focused on owning the reliability, scale, performance, and developer experience of core infrastructure systems. You will join an engineering team of roughly 15 people working on an AI/ML evaluation and reinforcement learning platform. Your work directly shapes how fast, reliable, and cost-effective the platform is to build on and operate.
Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
Design and improve backend and platform systems for scale, including capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected, debugged, and resolved quickly.
Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
Write clean, maintainable production code to automate systems, improve backend services, and build internal developer tooling.
2 to 4 years of experience owning production cloud infrastructure for a high-availability, user-facing platform, with responsibility for uptime, performance, deployment safety, and cost.
Deep hands-on experience with AWS and containerized systems; Terraform, Kubernetes/EKS, Docker, EC2, networking, load balancers, and secrets management strongly preferred.
Proven track record building or operating CI/CD, release automation, observability, alerting, and incident response systems.
Strong backend engineering judgment across service architecture, APIs, databases, async systems, queues, and production failure modes.
Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services.
Background operating infrastructure for AI/ML, data-heavy, marketplace, workflow, developer-tools, or enterprise platforms; generic DevOps or infrastructure-only experience is not sufficient.
Demonstrated focus on reducing cloud spend through better architecture, autoscaling, workload placement, caching, or cleanup systems.
Comfort writing production-quality code to automate and improve systems, not just configure them.
Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
On-site in Singapore. Candidates based in San Francisco or working remotely from other regions may also be considered depending on location; please review the role requirements carefully.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 219,313+ Jobs in Software Engineer
Answer easy questions
219,313+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”