Please mention DailyRemote when applying
See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewQuestions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will serve as the senior technical owner of the infrastructure, managing GKE clusters and Terraform while designing self-service platforms for developers. Additionally, you will define SLOs, mentor engineering staff, and drive continuous improvements in incident response and operational readiness.
InfiniteChoice is building its next-generation platform on Google Cloud. As our Staff Platform Engineer, you'll be the senior technical owner of that infrastructure: the GKE clusters, the Terraform that defines them, and the path traffic takes from the edge to our services. You'll also build what comes next: a self-service platform that lets our developers get infrastructure without filing a ticket, and the practices that keep our cloud footprint efficient as we grow
You'll report to the Director of Engineering, CloudOps, on a small team. Staff here means you set the technical direction for our infrastructure and raise the level of the engineers around you, while still writing a good share of the code yourself.
Take ownership of our GKE estate and Terraform, and get us to where every production infrastructure change goes through code review and Atlantis.
Bring structure to a GCP organization that grew quickly: project layout, group-based IAM, and cost visibility.
Ship the first self-service workflows for our developers, starting with the request types that fill most of our ticket queue today.
Turn our disaster recovery design into something we test on a schedule.
Participate in the on-call rotation and drive continuous improvements to incident response, operational readiness, and post-incident follow-through.
Design and run the platform our product teams deploy onto, including templates, CI/CD, and GKE Gateway API routing, so teams can ship without filing a ticket with us.
Define SLOs with the teams that own each service, and use them to decide where reliability work goes.
Make architecture decisions for our infrastructure and write them up as ADRs others can review.
Review infrastructure changes across engineering and mentor the engineers making them.
Plan capacity for clusters, databases, and caches, find and remove waste, and keep GCP spend in line with traffic as we grow.
10+ years of experience in platform, SRE or infrastructure engineering, including ownership of production systems serving millions of requests a day.
Deep hands-on experience running Kubernetes in production, ideally GKE.
Strong Terraform experience, including structuring modules and state across many environments and GCP projects.
Solid GCP experience across networking, IAM, and core services.
You've led a migration of production workloads between platforms without downtime.
You write production code in Go, Python, or a similar language, and you can read application code when debugging.
You explain technical tradeoffs clearly in writing, to engineers and to non-engineers.
Bachelor's degree in Computer Science, Engineering, or equivalent professional experience
Industry certifications (Google Cloud Professional, SRE or related certifications preferred)
Practical experience with AI coding tools, and a view on where they help in infrastructure work and where they don't
Cloudflare, including WAF and bot management
Redis at scale, including clustered deployments
OpenTelemetry and Prometheus
Building or running an on-call and incident program
Full autonomy to define processes, select technologies, and establish best practices
Direct impact on platform reliability serving millions of users
Opportunity to create lasting engineering culture and operational excellence
Remote-first culture with in-person meeting in Dallas, TX on need basis
Collaborative environment with smart, passionate engineers and cross-functional teams
Access to cutting-edge technologies and AI-driven development tools
Competitive compensation, equity participation, and comprehensive benefits
InfiniteChoice was founded to help people find the experiences they want simply and effortlessly. We leverage a new type of business model and platform that uniquely applies automation and technology to solve the challenges of scale and complexity in experience discovery.
Existing business and marketing technologies can no longer handle the demands of connecting millions of consumers with vast inventories of experiences across a fragmented, global marketplace of people, partners, and providers.
Our mission is to disrupt this status quo by creating seamless connections between consumers and experiences. We're just at the beginning of this journey, but our approach is working: we've helped over 10 Million+ customers, generating over $7 Billion in revenue for our brands and partners.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 214,723+ Jobs in Platform Engineer
Answer easy questions
214,723+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”