For Employers

MetaRouter

Principal Site Reliability Engineer

Posted 10 days ago
$180K - $250K per year
10+ years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

The Principal Site Reliability Engineer will define the reliability strategy, architecture, and observability standards for distributed systems across cloud and customer environments. They will lead major incidents, drive infrastructure automation, and mentor engineering teams to ensure high availability and compliance.

About The Role

As a Principal Site Reliability Engineer, you set the reliability strategy for the platform. You will define how we build, deploy, observe, and operate a distributed system that runs both in our own cloud and inside customer-controlled environments — and you will hold the organization to that standard.

We run dedicated, isolated environments per customer, which makes repeatability and automation the central engineering problem rather than an afterthought. Depth of judgment about reliability engineering matters far more here than experience with any particular cloud, orchestrator, or observability vendor.

This is an individual contributor role with organization-level influence.

Core Responsibilities

  • Own the reliability architecture of the platform: deployment topology, failure domains, capacity strategy, and the automation that makes environments reproducible.

  • Define service level objectives with product and engineering leadership, and drive the work needed to meet them.

  • Set the standard for observability — dashboards, logs, metrics, tracing, and alerting — so that issues are detected before customers report them.

  • Lead major incidents, run blameless postmortems, and make sure the corrective work actually lands.

  • Contribute to design and architecture across infrastructure and applications, with automation, performance, reliability, and security as first-class concerns.

  • Drive infrastructure lifecycle at scale: provisioning, upgrades, and decommissioning across many isolated environments.

  • Ensure infrastructure and applications meet or exceed enterprise compliance requirements, and design identity and access controls across platforms and services.

  • Partner with enterprise customers on custom infrastructure requirements, translating their constraints into repeatable patterns rather than one-off work.

  • Raise the bar through code and design review, and mentor SREs and product engineers on reliability practice.

  • Improve and maintain infrastructure and process documentation.

  • Participate in and help evolve the on-call rotation, including how the team balances operational load against project work.

Qualifications and Experience

  • 10+ years in infrastructure, SRE, or platform engineering, including deep experience operating large-scale distributed systems in production.

  • Expertise designing, analyzing, and troubleshooting distributed systems, with a track record of reliability decisions that held up under growth.

  • Deep experience with at least one major public cloud provider, and the ability to reason across providers rather than within one.

  • Strong command of container orchestration: cluster operation, workload scheduling, networking, and the failure modes that come with them.

  • Fluency with infrastructure as code, configuration management, and CI/CD pipeline design.

  • Strong scripting and automation ability, and comfort reading and debugging application code in the languages your services are written in.

  • Experience defining observability strategy — instrumentation, query languages, dashboards, and alert design that minimizes noise.

  • Demonstrated ability to influence without authority and align multiple teams behind a technical direction.

  • Experience operating under enterprise security and compliance frameworks.

  • Solid understanding of Unix/Linux operating systems and networking fundamentals.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

SME Software Developer/System Software

Full Time United States Software Development

Senior Backend Developer

Full Time Canada 150K - 170K per year Software Development

Data Analytics Consultant

Full Time Canada 125K - 135K per year Software Development

Senior Data Scientist – MMM ID71005

Full Time Colombia Software Development

Director, Presales Data Science

Full Time United States $132K - $207K per year Software Development

Machinery and Equipment - Field Engineer

Full Time United States $63000 - $105K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 216,664+ Jobs in Site Reliability Engineer

Answer easy questions

Answer easy questions

216,664+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified