For Employers

ECS Tech Inc

Cloud Site Reliability Engineer (SRE)

Posted 7 days ago
$130K - $180K per year
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review

Questions interviewers often ask for this role, with sample answers.

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

The role involves designing and building automated self-healing systems to ensure operational readiness and reliability for federal cloud platforms. You will define SLOs, enforce infrastructure-as-code standards, and lead incident management processes to minimize manual intervention.

Everforth ECS is seeking a Cloud Site Reliability Engineer (SRE) to work in our Arlington, VA office/remotely. 

 

Our Philosophy 
We believe the job of an SRE is to engineer the cloud to run itself. That means writing software and automation that lets systems detect and recover from failure on their own, rather than relying on someone to notice an alert and manually fix it. When something breaks, self-healing comes first, deep root-cause debugging happens after service is restored, not instead of it. We’re looking for someone who automates the operational task by default, not documents the runbook for doing it by hand. 
 
About the Role 
This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You’ll define what “reliable enough” looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation. 
 
Responsibilities 
Self-Healing Operations 

  • Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention 
  • Shift the team’s posture from “is it running, how do we fix it” to “how do we make it fix itself” 
  • Automate service restoration first; investigate root cause after 

Uptime Goals & Reliability 

  • Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them 
  • Use live metrics to decide what’s “reliable enough” and where to invest next 

Infrastructure 

  • Enforce infrastructure-as-code and configuration-as-code, no manual tech change 
  • Own Terraform standards and reusable modules adopted across programs 
  • Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling) 
  • Set CI/CD and pipeline-as-code standards, including progressive delivery 

Observability & Incidents 

  • Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct 
  • Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation 

Collaboration & Leadership 

  • Partner with development and contractor teams leads to embed reliability and automation across the software 
  • Mentor engineers toward this same automation-first philosophy 
  • Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation 

 

 

Salary Range: $130,000 - $180,000 

General Description of Benefit

Qualifications
  • Bachelor’s degree in Computer Science, Information Technology, or related field (or equivalent practical experience) 
  • 5+ years of SRE experience (or equivalent), with demonstrated technical leadership 
  • 10 years of general work experience
  • Track record building self-healing/auto-remediating systems, not just dashboards 
  • Jenkins experience 
  • Expert AWS knowledge, GovCloud experience strongly preferred 
  • Deep Kubernetes and Terraform expertise at production scale 
  • Strong software engineering background (Python and/or Go) 
  • Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc. 
  • Proven incident command and postmortem experience 
  • Strong communication skills across technical and federal leadership audiences 
  • Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable) 
  • US Citizenship 

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Senior Systems Engineer

Full Time United States $144 per hour Software Development

Principal Solutions Architect

Full Time United States Software Development

Continuous Improvement & Analytics Manager

Full Time United States Software Development

AP/AR Lead

Full Time Philippines Software Development

Senior Linux Systems Administrator (Work from Home)

Full Time Philippines Software Development

Backend Developer (Node.js/Nest.js)

Full Time Armenia, Cyprus, Georgia +2 more Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 216,786+ Jobs in Site Reliability Engineer

Answer easy questions

Answer easy questions

216,786+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified