For Employers

HostPapa, Inc.

Site Reliability Engineer

Posted 4 hours ago
2-5 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review

Questions interviewers often ask for this role, with sample answers.

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

The Site Reliability Engineer will ensure the reliability, scalability, and observability of multi-tenant SaaS platforms through monitoring and incident response. They will also collaborate with engineering teams to automate operational tasks and design resilient, fault-tolerant system architectures.

Position Summary: 

At managed.com, we're a remote-first, global team. We're also AI-first, so finding practical ways to use AI is part of every role here.

This role focuses on CloudBlue, a managed.com business that powers cloud commerce for many of the world’s largest service providers, including major Telcos, distributors, and MSPs. CloudBlue enables partners to monetize and manage cloud services and subscriptions at scale, combining the agility of a high-growth business with the backing of a global organization.

 

As the Site Reliability Engineer, you will help ensure the reliability, scalability, and observability of CloudBlue’s multi-tenant SaaS platforms used by service providers worldwide. You will focus on improving system stability and performance through monitoring, high availability, and incident response, while working closely with DevOps, Platform, and Engineering teams to build and operate resilient production systems.

What you’ll do

  • Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services to ensure reliability and performance
  • Influence system architecture with a strong focus on reliability, scalability, and operability, designing systems for fault tolerance, graceful degradation, and self-healing
  • Reduce operational toil by identifying opportunities for automation and process improvement
  • Design and operate CloudBlue’s observability stack across metrics, logs, and traces using tools such as Datadog, Grafana, and Elastic Stack
  • Develop actionable alerting strategies and dashboards that provide clear insight into platform and business health
  • Design and maintain high-availability architectures, implementing redundancy, failover, and disaster recovery strategies across regions and availability zones
  • Conduct capacity planning, load testing, and performance optimization to ensure platform stability and scalability
  • Act as a senior responder during production incidents, leading incident coordination, communication, and service restoration
  • Own blameless postmortems and drive improvements that reduce incident frequency, MTTR, and customer impact
  • Improve reliability of Kubernetes-based platforms through health checks, autoscaling strategies, rollout safety, and resilience testing
  • Partner with engineering and DevOps teams to improve deployment safety, rollback strategies, and platform reliability
  • Maintain runbooks and operational documentation, and promote SRE best practices across engineering teams
  • Support other tasks or projects as assigned to meet team and business needs

About you

  • 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer, with strong ownership of production systems
  • Proven experience operating highly available, enterprise-grade, multi-tenant SaaS platforms
  • Hands-on experience with observability and monitoring tools such as Datadog, Grafana, and Elasticsearch/Kibana
  • Solid understanding of Linux, networking, and distributed systems fundamentals
  • Experience working with containerized environments such as Docker and Kubernetes
  • Strong scripting and automation skills using Python and/or Bash
  • Experience participating in on-call rotations and incident response in production environments
  • Strong written and spoken English
  • Experience defining SLIs/SLOs and managing error budgets at scale will be considered a plus
  • Exposure to hyperscale or service-provider-grade platforms is an advantage
  • Cloud experience, preferably with Azure; experience with AWS and/or GCP will also be valued
  • Experience working with hybrid or on-premises integrations is beneficial
  • Familiarity with chaos engineering and resilience testing will be considered an asset

What we offer:

  • Remote-first work, with team members across many countries
  • Competitive pay, benchmarked against the markets we hire in
  • Room to grow, with real ownership of your work and support for professional development
  • Flexibility in how and when you work, within the core hours your team needs
  • Access to the AI tools you need to do your work, and the room to experiment with them
 

About us:

Managed.com is the company behind a family of specialist brands spanning business technology, infrastructure, and channel commerce. These include HostPapa, ColoCrossing, Hostopia, CloudBlue, and LogoMaker. Each brand serves its market through its own products and customer experience. Together, we are bringing those strengths closer to create more connected experiences and help businesses build, operate, and grow in the AI era.

Our culture is built on trust, respect, and doing work that matters to our customers. As our brands come together, there's real room to grow into new roles and new areas of the business.

Managed.com is an equal-opportunity employer. We value the range of backgrounds, perspectives, and experiences across our team, and we hire and promote based on skills and contribution.

We're committed to providing accommodations for people with disabilities at every stage of the hiring process. If you need an accommodation, let us know, and we'll work with you to meet your needs.

It is anticipated that this position will be performed outside of Ontario.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Assistant Vice President - Data Analytics – Retirement Division - Remote

Full Time United States $188K - $314K per year Software Development

Sr AI/ML Engineer (LATAM Remote)

Full Time Mexico, United States Software Development

Solution Architect

Full Time Portugal Software Development

Infrastructure Reliability Engineer

Full Time United States $100K - $150K per year Software Development

AI Operations Engineer

Full Time United States $150K - $165K per year Software Development

AI Data Platform Engineer

Full Time United States $100K - $150K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 220,094+ Jobs in Site Reliability Engineer

Answer easy questions

Answer easy questions

220,094+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified