Please mention DailyRemote when applying
See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewQuestions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
The Principal Site Reliability Engineer will define the reliability strategy, architecture, and observability standards for distributed systems across cloud and customer environments. They will lead major incidents, drive infrastructure automation, and mentor engineering teams to ensure high availability and compliance.
As a Principal Site Reliability Engineer, you set the reliability strategy for the platform. You will define how we build, deploy, observe, and operate a distributed system that runs both in our own cloud and inside customer-controlled environments — and you will hold the organization to that standard.
We run dedicated, isolated environments per customer, which makes repeatability and automation the central engineering problem rather than an afterthought. Depth of judgment about reliability engineering matters far more here than experience with any particular cloud, orchestrator, or observability vendor.
This is an individual contributor role with organization-level influence.
Own the reliability architecture of the platform: deployment topology, failure domains, capacity strategy, and the automation that makes environments reproducible.
Define service level objectives with product and engineering leadership, and drive the work needed to meet them.
Set the standard for observability — dashboards, logs, metrics, tracing, and alerting — so that issues are detected before customers report them.
Lead major incidents, run blameless postmortems, and make sure the corrective work actually lands.
Contribute to design and architecture across infrastructure and applications, with automation, performance, reliability, and security as first-class concerns.
Drive infrastructure lifecycle at scale: provisioning, upgrades, and decommissioning across many isolated environments.
Ensure infrastructure and applications meet or exceed enterprise compliance requirements, and design identity and access controls across platforms and services.
Partner with enterprise customers on custom infrastructure requirements, translating their constraints into repeatable patterns rather than one-off work.
Raise the bar through code and design review, and mentor SREs and product engineers on reliability practice.
Improve and maintain infrastructure and process documentation.
Participate in and help evolve the on-call rotation, including how the team balances operational load against project work.
10+ years in infrastructure, SRE, or platform engineering, including deep experience operating large-scale distributed systems in production.
Expertise designing, analyzing, and troubleshooting distributed systems, with a track record of reliability decisions that held up under growth.
Deep experience with at least one major public cloud provider, and the ability to reason across providers rather than within one.
Strong command of container orchestration: cluster operation, workload scheduling, networking, and the failure modes that come with them.
Fluency with infrastructure as code, configuration management, and CI/CD pipeline design.
Strong scripting and automation ability, and comfort reading and debugging application code in the languages your services are written in.
Experience defining observability strategy — instrumentation, query languages, dashboards, and alert design that minimizes noise.
Demonstrated ability to influence without authority and align multiple teams behind a technical direction.
Experience operating under enterprise security and compliance frameworks.
Solid understanding of Unix/Linux operating systems and networking fundamentals.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 216,664+ Jobs in Site Reliability Engineer
Answer easy questions
216,664+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”