For Employers

STN Inc

Site Reliability Engineer

Posted 2 months ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

The SRE owns reliability, observability, and incident response for the GPUaaS platform. Key duties include defining SLOs, building the observability stack, and leading major incident resolution.

Site Reliability Engineer

Platform and software · shared across customers

Reports to: Director, Site Reliability

Location: Remote (US)

Department: Cloud Platform Engineering / SRE/Reliability

Position summary

The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.

Key responsibilities

  • Define and operate Service Level Objectives (SLOs) aligned with customer SLAs

  • Build and maintain the observability stack including metrics, logs, traces, and alerting

  • Lead incident response and chair post-incident reviews

  • Drive automation to reduce toil and improve mean-time-to-recover (MTTR)

  • Author and maintain operational runbooks alongside the NOC

  • Manage on-call rotation, escalation paths, and incident-management tooling

  • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering

  • Drive chaos engineering, game days, and reliability testing programs

  • Produce SLA performance reports in coordination with the SLA Manager

  • Mentor junior engineers and contribute to engineering culture

Required qualifications

  • 5+ years in SRE, DevOps, or production engineering roles

  • Strong programming skills in Go, Python, or both

  • Hands-on experience operating Kubernetes-based platforms at scale

  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)

  • Strong incident management experience including major-incident command

Preferred qualifications

  • GPU or HPC platform operational experience

  • Familiarity with SLA-driven customer environments and credit calculations

  • Experience with chaos engineering tools (Gremlin, Litmus, or similar)

  • Published SRE content or contributions

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

R&D Engineer 5, Software

Full Time Finland Software Development

Vice President of Solutions Marketing & Global Product Advocacy | SaaS | Remote | 50% to 75% Travel

Full Time United States Software Development

LINUX System Administrator

Full Time India Software Development

Senior/Principal FPGA Design Engineer

Full Time Canada, United States Software Development

ServiceNow Platform Architect

Full Time Poland Software Development

RF/MW Solutions Engineer – Aerospace & Defense

Full Time United Kingdom Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Site Reliability Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified