For Employers

Glint Tech Solutions LLC

Lead Site Reliability Engineer

Posted 2 months ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

The Lead Site Reliability Engineer will design and implement SRE practices to ensure the reliability, scalability, and performance of critical banking platforms. They will lead incident response, drive automation, and mentor engineering teams on observability and cloud-native architecture.

Job Title: Lead Site Reliability Engineer

Location: Remote within the USA, or onsite in Buffalo, NY / Wilmington, DE (client preference for candidates near these areas). New hires are required to work onsite at the client's office for the first 2–3 weeks (treated as a business trip; travel expenses covered by the company).

Company Overview
Glint Tech Solutions is a women-owned, global IT staffing and recruiting firm serving enterprise clients across the USA and Canada.

Project Description
A leading financial services client is seeking a Lead Site Reliability Engineer responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. This senior individual contributor will design, implement, and improve SRE practices across the software development lifecycle, working closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management, while coaching and influencing others.

Key Responsibilities

  • Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise SRE best practices
  • Define, implement, and monitor SLOs, SLIs, and error budgets for critical business services
  • Develop observability strategies using Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting
  • Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints
  • Lead incident response for high-severity production events and facilitate Root Cause Analysis (RCA)
  • Drive operational excellence through automation of deployments, recovery procedures, and reliability controls
  • Design and execute automated regression testing strategies to validate stability and performance
  • Create and maintain Infrastructure as Code (IaC) solutions using Terraform
  • Support and optimize Microsoft Azure environments, including App Services, scaling, and deployment automation
  • Utilize Azure Monitor, Application Insights, and Log Analytics to improve platform visibility
  • Drive performance testing, resiliency testing, and disaster recovery preparedness
  • Lead capacity planning, performance tuning, and workload optimization
  • Develop operational runbooks, incident playbooks, and standard operating procedures
  • Mentor engineers on observability, cloud engineering, automation, and SRE principles
  • Adhere to Company risk and regulatory standards, policies, and controls

Mandatory Skills

  • Strong hands-on experience with Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics collection, and centralized logging
  • Proven experience designing and executing automated regression testing frameworks
  • Strong proficiency in Infrastructure as Code (IaC) using Terraform
  • Experience with CI/CD pipelines, deployment automation, and operational tooling
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting
  • Strong understanding of application performance management, distributed systems, and cloud-native architectures
  • Strong experience with Microsoft Azure (App Services, Resource Groups, networking, scaling, deployment/release management)
  • Experience with Azure Monitor, Application Insights, Log Analytics, and Azure dashboards/alerting
  • Experience supporting cloud-native and hybrid infrastructure environments
  • Demonstrated experience implementing SRE practices — SLOs, SLIs, error budgets, incident/problem management, RCA, reliability automation
  • Ability to improve system reliability through performance tuning, capacity planning, and observability-driven insights
  • Experience developing automated recovery mechanisms and self-healing solutions
  • Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures

Nice-to-Have Skills

  • Experience supporting large-scale enterprise applications in regulated environments
  • Experience working in Agile and DevOps operating models
  • Ability to work autonomously and lead complex reliability initiatives
  • Experience partnering with architecture, infrastructure, cybersecurity, and application development teams
  • Scripting/automation experience with PowerShell, Python, or Bash
  • Industry certifications in Azure, Terraform, Cloud Engineering, or Site Reliability Engineering
  • Proven experience leading major incident response and post-incident improvement efforts

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Senior Analyst, Growth & Marketing Analytics

Full Time United States $112K - $160K per year Software Development

Application Engineer, Data Centers

Full Time United States $77600 - $124K per year Software Development

Data Analyst IV

Full Time United States $118K - $155K per year Software Development

IS Data Warehouse Architect III

Full Time United States $126K - $154K per year Software Development

Senior Integration Developer

Full Time United States Software Development

Salesforce Technical Architect - Remote

Full Time United States $80 - $100 per hour Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 218,390+ Jobs in Site Reliability Engineer

Answer easy questions

Answer easy questions

218,390+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified