For Employers

Simple Life App

Senior Site Reliability Engineer

Posted a day ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

Operate and improve AWS infrastructure, the Kubernetes platform, infrastructure-as-code, CI/CD, observability, and the on-call rotation. Investigate production incidents to identify root causes and build automation and internal tools, primarily in Go.

Simple Life is the #1 AI-powered health coaching app for adults who want to lose weight and enjoy a healthier lifestyle—without the stress or extremes. Our mission is to empower people to feel their best every day. By challenging traditional, restrictive approaches, Simple offers a more sustainable method grounded in ease, personalization, and real-life support.

Simple has had over 17 million downloads and more than 300,000 5-star reviews, having helped millions lose weight successfully and sustainably. Simple has earned recognition as Best Virtual Coach and one of the Top 100 AI companies — all thanks to a dedicated global team driving real impact.

With SIMPLE as a partner in their pocket, users feel cared for and empowered to embrace — and stick to — new healthy habits. To learn more, visit simple.life.

Simple is looking for a Senior SRE to join our Platform team responsible for the AWS infrastructure, the Kubernetes platform and the internal tooling that the rest of engineering relies on.

Push the pace of innovation and build a future of a healthier world with us!

About the role:

This is an operations-led position. You will be working day to day on AWS, our infrastructure-as-code, our CI/CD setup, observability, and the on-call rotation.

A meaningful part of the work is automation: when we find ourselves doing the same thing twice, we usually invest in tooling rather than writing another runbook. Most of that tooling is written in Go.

To give a sense of the environment: infrastructure is defined with Terraform and Terramate, with Atlantis running plan and apply on pull requests. Workloads run on EKS with Karpenter, Cilium and Istio, deployed through ArgoCD. Observability is built on Grafana, Loki, Tempo, and Prometheus compatible metrics.

We’re looking for:

  • The most important quality for this role is how you handle problems that are not yet understood.

  • Production incidents rarely present cleanly: logs can be incomplete, metrics can mislead, and the first plausible theory is often wrong.

  • The right candidate stays focused under that kind of pressure, works through ambiguity in a structured way, and arrives at a real root cause rather than a convenient one.

  • Strong investigation, debugging, and problem-solving instincts are essential.

  • We also expect candidates to learn quickly. The stack and the company both move, and you will regularly be the first person on the team to take on something new.

  • Several years of hands-on experience operating production systems on AWS and Kubernetes, including genuine on-call ownership.

  • A solid working knowledge of AWS fundamentals, including VPC, IAM, EKS, and RDS.

  • Practical experience with Terraform and a GitOps-style delivery workflow (ArgoCD, Atlantis, Flux, or similar).

  • Comfort writing code, with some prior experience in Go or a willingness to pick it up (writing small services and tools is a regular part of the work).

  • Strong written and spoken English, and the communication skills to drive design discussions across engineering, product and security.

  • We're flexible on some locations as long as your working hours overlap with European time zones

Perks and Benefits:

  • Open-minded teams, a welcoming and inclusive company culture, plus the opportunity to make a real difference with a game-changing health tech product.

  • A competitive salary package based on your unique expertise, skillset, and impact on the product plus stock options.

  • In-office, remote and hybrid work opportunities.

  • The equipment whatever you need to be happy and productive.

  • A premium SIMPLE subscription.

  • 21 days annual leave, plus bank holidays

  • Flexible hours. We focus on your results, not how long you spend at your desk.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Senior Data Engineer (all genders)

Full Time, Part Time Germany Software Development

Developer

Full Time Portugal Software Development

Site Reliability Engineer

Full Time Argentina, Brazil, Costa Rica +1 more Software Development

Batch Developer (UNIX/C/SQL)

Full Time United States $60000 - $80000 per year Software Development

Outside Sales Engineer - OEM (South)

Full Time United States $83500 - $130K per year Software Development

Outside Sales Engineer - OEM (Midwest US)

Full Time United States $83500 - $130K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 216,738+ Jobs in Site Reliability Engineer

Answer easy questions

Answer easy questions

216,738+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified