For Employers

Encora

Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)

Posted 3 hours ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

Collaborate with development teams to design and implement monitoring, alerting, and APM instrumentation across services. Lead the optimization of observability solutions and define service level objectives to ensure application reliability and performance.

Job Title: Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)
Key Skills: Azure IaaS, Azure Monitor, Application Insights, New Relic, Databricks, DBT, SQL, Azure DevOps, GitHub Actions, Site Reliability Engineering (SRE), Application Performance Monitoring (APM), Log Analytics (KQL), Observability, Distributed Tracing, Structured Logging
Experience: 5+ years of experience in Site Reliability Engineering, Cloud Operations, or related roles. Mandatory minimum of 1 year of hands-on experience with DBT, Databricks, and SQL.
Location: Legal residents of Peru, Colombia, Bolivia, Costa Rica, Mexico, and Brazil.
Work Mode: Remote
 
At Coforge, we are looking for a Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure) (#21013-18-1) with the following profile.

Main Responsibilities

  • Collaborate with development teams to design and implement monitoring, alerting, dashboards, and APM instrumentation across applications and services.
  • Lead the implementation, configuration, and optimization of Application Performance Monitoring (APM) solutions.
  • Apply observability best practices using tools such as Azure Monitor, Application Insights, New Relic, and Log Analytics (KQL).
  • Enable code-level instrumentation, distributed tracing, and structured logging to improve application visibility and reliability.
  • Design and maintain application-level monitoring dashboards and operational health metrics.
  • Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and effective alerting strategies based on latency, error rates, traffic, and resource saturation.
  • Continuously improve monitoring and alerting mechanisms through production insights and incident learnings.
  • Participate in production readiness reviews, identifying operational risks, observability gaps, and potential failure scenarios before deployment.
  • Support incident analysis and post-incident improvements through enhanced telemetry and monitoring practices.
  • Partner with engineering teams to ensure applications are reliable, scalable, and production-ready.

Mandatory Requirements

  • Strong experience supporting and operating applications in Microsoft Azure IaaS environments.
  • Hands-on experience with application observability, monitoring, and reliability engineering practices.
  • Mandatory experience with DBT, Databricks, and SQL (minimum 1 year of experience).
  • Experience implementing and managing APM solutions such as Application Insights, New Relic, or similar platforms.
  • Experience designing dashboards and monitoring solutions using Azure Monitor, Application Insights, and Log Analytics (KQL).
  • Familiarity with CI/CD environments including Azure DevOps and GitHub Actions.
  • Solid understanding of cloud-native architectures and distributed application systems.
  • Practical SRE mindset with experience in incident analysis, root cause investigation, and proactive problem prevention.
  • Strong verbal and written English communication skills, with the ability to collaborate effectively with global teams.

Preferred Requirements

  • Experience with scripting and automation using PowerShell and/or Bash.
  • Knowledge of scalability, availability, and resilience patterns in modern cloud environments.
  • Experience driving production readiness and operational excellence initiatives.
  • Exposure to reliability engineering best practices in enterprise-scale environments.
 
Posted on: 28-09-2026
At Coforge, we hire professionals solely based on their skills and qualifications and do not discriminate on the basis of age, disability, religion, gender, sexual orientation, socioeconomic status, or nationality.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Principal Software Engineer

Full Time United States $142K - $304K per year Software Development

SENIOR SALESFORCE BUSINESS ANALYST / PRODUCT OWNER

Full Time India 2500K - 3500K per year Software Development

Google AI Engineer

Full Time Mexico Software Development

IT Systems Engineer

Full Time Canada Software Development

Senior Software Engineer, Frontend

Full Time United States Software Development

Sales Specialist HPC/AI Products (Southeastern US)

Full Time Finland $216K - $507K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 219,670+ Jobs in Site Reliability Engineer

Answer easy questions

Answer easy questions

219,670+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified