[8SN] Senior Site Reliability / Production Support Engineer

 Posted 15 hours ago
  
 Canada
  
5-10 years experience
Apply Now

Please mention DailyRemote when applying

AI Summary

The role involves supporting the deployment, operations, and reliability of production services running on Kubernetes. You will monitor service health, investigate incidents, and collaborate with engineering teams to improve operational efficiency.

Company Description

We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
 

About the Client

Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies.

You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production.

#LI-DNI

Job Description

About the Role

We are looking for a Senior Site Reliability / Production Support Engineer to support the deployment, operations, and ongoing reliability of a production UI service running on Kubernetes.

This role is focused on maintaining highly available cloud-native applications, troubleshooting production issues, and improving operational excellence. You will work closely with engineering teams to monitor service health, investigate incidents, and ensure reliable service delivery.

While this role supports a UI-based service, it is not a frontend development position. Basic knowledge of Web Components is sufficient to perform first-level debugging when necessary.
 

What You'll Do

  • Support deployment, operations, and ongoing maintenance of a production service running on Kubernetes.
  • Monitor application health, availability, and performance.
  • Investigate and resolve production incidents using logs, monitoring, and debugging tools.
  • Perform log analysis using Splunk to identify root causes and troubleshoot service issues.
  • Collaborate with software engineers to improve service reliability and operational efficiency.
  • Participate in incident response and production support activities.
  • Assist with first-level debugging of UI-related issues involving Web Components.
  • Contribute to continuous improvements in automation, monitoring, and operational processes.
  • Support CI/CD pipelines and cloud-native deployment practices.

Qualifications

Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Operations.
  • Strong hands-on experience with Kubernetes in production environments.
  • Experience supporting cloud-native applications.
  • Experience monitoring production systems and troubleshooting complex incidents.
  • Strong knowledge of Splunk for log analysis and debugging.
  • Experience working in Linux environments.
  • Understanding of networking fundamentals and distributed systems.
  • Experience collaborating with software engineering teams to resolve production issues.
  • Strong troubleshooting and root cause analysis skills.
  • Excellent written and spoken English (B2+).

Additional Information

Preferred Qualifications

  • Experience with CI/CD pipelines.
  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Familiarity with container technologies such as Docker.
  • Exposure to observability tools (Prometheus, Grafana, OpenTelemetry, etc.).
  • Basic understanding of Web Components and frontend architecture.
  • Experience supporting high-availability enterprise SaaS platforms.
  • Knowledge of infrastructure automation or Infrastructure as Code (Terraform, Helm, Ansible, etc.) is a plus.

What We Offer

  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements and collaborative team environment
  • Impact on organization-wide digital transformation initiatives

Similar Jobs

See all Remote Software Development jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Support Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified