Please mention DailyRemote when applying
We're a fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. Our engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues — and we're looking for a Platform Engineer to own the reliability, scale, performance, and developer experience of our core infrastructure.
This is a backend-architecture-heavy role with high ownership. Your work will directly determine how fast, reliable, and cost-effective our platform is to build on and run.
Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected, debugged, and resolved quickly.
Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
Write clean, maintainable production code to automate systems, improve backend services, and create internal developer tooling.
Required:
2–4 years of experience owning production cloud infrastructure for a high-availability, user-facing platform, with accountability for uptime, performance, deployment safety, and cost.
Deep hands-on experience with AWS and containerized systems; strong familiarity with Terraform, Kubernetes/EKS, Docker, EC2, load balancers, networking, and secrets management.
Track record of building or operating CI/CD, release automation, observability, alerting, and incident response systems.
Strong backend engineering judgment — able to reason about service architecture, APIs, databases, async systems, queues, scaling limits, and production failure modes.
Ability to write production-quality code to automate infrastructure, improve backend services, and build internal tooling.
Nice to Have:
Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services.
Background operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms.
Demonstrated focus on reducing cloud spend through better architecture, autoscaling, workload placement, caching, or cleanup systems.
Experience building internal platforms or developer tools that improve engineering productivity without hiding complexity.
We prioritize technical aptitude, ownership, and learning potential over years of experience.
San Francisco, CA (on-site): US-based candidates must be located in San Francisco.
Singapore (on-site): Southeast Asia-based candidates must be located in Singapore.
Fully remote (contractor): Candidates based elsewhere — particularly in Europe — may be considered as fully remote independent contractors.
Visa sponsorship is available.
Salary: $150,000 – $250,000 USD annually (for full-time roles)
Opportunity to have significant ownership and direct impact at an early-stage, well-funded AI infrastructure company.
Work alongside a world-class technical team building foundational infrastructure for AI alignment and post-training data.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Platform Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!