The Senior Site Reliability Engineer ensures the reliability, performance, and scalability of critical systems across AWS and on-premises environments by improving deployment processes and increasing automation. This role involves transitioning applications to containerized deployments, managing incident response, and enhancing system monitoring and observability.
Senior Site Reliability Engineer
Location: Remote
Compensation: $160,000 - 180,000 per year, depending on experience and qualifications. Employment Type: Full-Time
What you can expect as the Senior Site Reliability Engineer at Fortress… The Senior Site Reliability Engineer is responsible for ensuring the reliability, performance, and scalability of critical systems and services across AWS and on-premises environments. This role focuses on improving deployment processes, strengthening CI/CD pipelines, increasing automation, and enhancing system monitoring and observability. The position also leads key efforts such as transitioning applications to containerized deployments, implementing advanced deployment strategies, and improving shared infrastructure for developers. Candidates in this role will support production environments, manage incident response, and contribute to continuous improvement of Fortress’ operational reliability.
Responsibilities Include
Transition applications from traditional (non-containerized) Ansible deployments to containerized, orchestrated deployments in AWS and on-premises environments
Build upon current CI/CD efforts to support different deployment strategies (Blue/Green, Canary, etc.)
Support Development/QA/UAT efforts by building an environment for anyone in the company to test drive a release anywhere in the development lifecycle
Improve common infrastructure for developers, such as CI/CD pipelines, log/application monitoring, cluster management, and configuration management
Migrate application secrets and configuration from Ansible Vault to Hashicorp Vault
Automate infrastructure provisioning during deployment
Handle code deployments in all environments (cloud, on-premises)
Implement tools to monitor and alert with respect to service level metrics and objectives
Provide technical guidance and educate team members and coworkers on development and operations
Monitor relevant systems for availability and performance
Available to support daytime and after business hours release activities
Minimum Qualifications
5-8 years hands-on experience in a SRE/DevOps role supporting production systems
Production experience supporting Linux-based infrastructure and administering services on AWS (RDS, VPC, ECR, CloudWatch, Cloud Formation, Lambda, API Gateway) and on-premises
Demonstrable experience with deployment technologies such as Kubernetes, Ansible, Jenkins and Terraform (or similar technologies)
Excellent documentation skills so anyone on the team can come up-to-speed on changes quickly
Excellent written and verbal communication skills
Experience implementing rolling upgrades (canary, blue/green, etc.)
Strong scripting and tooling skillset (Bash, Python, JS, etc.)
Self-motivated, resourceful and a persistent problem-solving aptitude with advanced time management skills
Ability to independently use and refine prompts to enhance the quality, efficiency, and insight of regular work processes
Must be willing to participate in technical interviews and technical questions, which may be recorded or transcribed for evaluation purposes (required)
Preferred Skills
SOC, NERC, and other compliance standards knowledge a plus
Presentation skills and the ability to communicate with a variety of technical and non-technical audiences.
Strong Linux system administration skills
Strong scripting and tooling skillset (Bash, Python, JS, etc.)
Adept at creating comprehensive documentation
AWS administration
Education
Bachelor's Degree in Information Technology, Computer Science, or a related discipline required.
Master’s Degree in related field (preferred)
Employee Benefits
Remote and Hybrid working environment
Competitive pay structure
Medical, dental, vision plans with employees covered up to 90% with highly progressive options for dependents and families
Company paid life, short- and long-term disability insurance
Employee Assistance Program
401(k) match
Flexible Paid Time Off
Parental Leave
Employment Perks
We provide each employee with professional growth opportunities through succession planning, up-skilling, and certifications
Tuition and certification reimbursement
Employee Referral Programs
Company Sponsored Events
Fortress is proud to be an Equal Opportunity Employer. All employees and applicants will receive consideration for employment without regard to age, color, disability, gender, national origin, race, religion, sexual orientation, gender identity, protected veteran status, or any other classification protected by federal, state, or local law. Fortress Information Security takes part in the E-Verify process for all new hires.
For positions located in the US, the following conditions apply. If you are made a conditional offer of employment, you will have to undergo a drug test. ADA Disclaimer: In developing this job description care was taken to include all competencies needed to successfully perform in this position. However, for Americans with Disabilities Act (ADA) purposes, the essential functions of the job may or may not have been described for purposes of ADA reasonable accommodation. All reasonable accommodation requests will be reviewed and evaluated on a case-by-case basis.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”