For Employers

Halo Media

Site Reliability Engineer (SRE)

Posted an hour ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

The SRE will lead large-scale operating system modernization projects and drive infrastructure automation to improve system reliability. They will also provide operational support, manage incident response, and partner with engineering teams on cloud migration initiatives.

About the Role:

We are looking for a Senior Site Reliability Engineer (SRE) to help modernize large-scale infrastructure and improve the reliability, scalability, and operational excellence of critical production systems. In this role, you will lead OS modernization initiatives, drive infrastructure automation, strengthen observability, and partner closely with engineering teams to support cloud migration efforts.

The ideal candidate has a strong background in Linux systems, Python, automation, CI/CD, and production operations, with a passion for building resilient platforms and solving complex infrastructure challenges.

Responsibilities:

  • Lead large-scale operating system modernization projects, including migrations from RHEL7 to EL8/9 across approximately 1,700 systems and virtual machines.

  • Drive infrastructure and packaging migrations, including Chef to CINC and yinst to RPM.

  • Build, maintain, and configure RPM packages to support modern infrastructure deployments.

  • Develop automated operational runbooks and infrastructure automation to improve efficiency and reliability.

  • Harden CI/CD pipelines, rollout/rollback mechanisms, and deployment processes for infrastructure modernization.

  • Strengthen observability by onboarding services to modern monitoring and logging platforms.

  • Triage, investigate, and resolve complex production incidents and critical software bugs.

  • Provide Tier-2 operational support in a follow-the-sun model alongside Site Reliability Engineering and Cloud Infrastructure teams.

  • Support incident response, troubleshooting, and break/fix activities across distributed production environments.

  • Partner with software engineering teams during cloud migration initiatives, providing operational guidance and technical support.

  • Automate repetitive operational tasks and maintain comprehensive technical documentation.

  • Drive reliability, operational excellence, and continuous improvement across production systems.

Required Qualifications:

  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Software Engineering with a strong infrastructure focus.

  • Strong hands-on experience with Python.

  • Proven experience leading Linux operating system modernization and infrastructure migration projects.

  • Experience with RHEL, Linux administration, and package management.

  • Hands-on experience building and maintaining RPM packages.

  • Strong experience with infrastructure automation and configuration management.

  • Experience designing, maintaining, and improving CI/CD pipelines.

  • Strong troubleshooting and incident management skills in large-scale production environments.

  • Experience supporting distributed systems with a strong focus on reliability and availability.

  • Excellent scripting, automation, and problem-solving skills.

Benefits:

  • 100% Remote.

  • Salary in USD.

  • International and collaborative environment.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Senior Full Stack Engineer, Revenue Engine

Full Time Software Development

Senior Business Data Analyst - Remote in PA

Full Time $72800 - $130K per year Software Development

AI Tech Lead - GraphAware Hume

Full Time Software Development

Senior Software Engineer - GraphAware Hume

Full Time Software Development

Senior Platform Engineer, Infrastructure

Full Time $186K - $255K per year Software Development

SAP SD Consultant/Developer – AI Integration (f/m/d)

Full Time €60000 - €90000 per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Site Reliability Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified