Senior Site Reliability Engineer

 Posted an hour ago
     
2-5 years experience
Apply Now

Please mention DailyRemote when applying

AI Summary

You will design and evolve scalable GCP infrastructure while building internal tooling to improve developer productivity. Additionally, you will champion reliability practices, manage incident responses, and optimize cloud infrastructure costs.

About the Role

We are a well-funded AI/ML company operating at the intersection of geospatial intelligence and climate technology. Our engineering team builds products that rely on robust, scalable cloud infrastructure — and we're looking for a Senior Site Reliability Engineer to help us evolve that foundation.

In this role, you'll own and advance our GCP infrastructure and DevOps practices, driving improvements in incident management, SLOs, error budgets, and observability. You'll partner closely with Product & Engineering teams to use DORA metrics as a lever for continuous improvement, while also helping the broader organization understand and optimize cloud costs.

What You'll Do

  • Design and evolve our cloud infrastructure on GCP for scale and resilience.

  • Build internal tooling and automation that promote team autonomy and developer productivity.

  • Advance our observability platform — metrics, logging, tracing, and alerting — to reduce mean time to recovery.

  • Build visibility into infrastructure costs and drive optimization initiatives.

  • Champion reliability best practices across engineering, including SLOs/SLIs, error budgets, and post-incident reviews.

  • Participate in on-call rotation and lead incident management efforts.

What We're Looking For

Required:

  • 3+ years of Site Reliability Engineering or production SRE experience.

  • Proficiency with Google Cloud Platform (GCP), including cost optimization and governance.

  • Hands-on experience with Kubernetes for cluster and workload management.

  • Infrastructure as Code experience using tools such as Terraform or Deployment Manager.

  • Scripting and automation skills in Python, Bash, or Go.

  • Strong observability stack experience: Prometheus, Grafana, OpenTelemetry, logging, and tracing.

  • Demonstrated ability to define and implement SLOs, SLIs, and error budgets.

  • Experience with incident management, post-incident reviews, and on-call rotation.

Nice to Have:

  • Experience designing and evolving cloud infrastructure at scale.

  • Familiarity with DORA metrics and using them to drive engineering effectiveness.

Please note: Visa sponsorship is not available for this role.

Location

This is a fully remote role open to candidates based in the EU, UK, or North America (including Canada, Denmark, Estonia, France, Netherlands, Portugal, Sweden, Switzerland, and the United Kingdom).

Compensation & Benefits

Compensation details were not specified for this role. We are happy to discuss salary expectations during the interview process.

Similar Jobs

See all Remote Software Development jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Site Reliability Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified