Please mention DailyRemote when applying
Job Title: Site Reliability Engineer (SRE)
Key Skills: Kubernetes, AWS/Azure/GCP, Terraform, Python, Observability, CI/CD
Experience: +6 YOE.
Location: Costa Rica, Peru, Colombia, and Bolivia.
Mode: Remote.
We at Coforge are hiring Site Reliability Engineer (SRE) (#22323) with the following skill set.
Key Responsibilities
· Design, build, and operate scalable and highly available cloud platforms.
· Ensure reliability, performance, and stability of distributed production systems.
· Implement and maintain Infrastructure as Code using Terraform or similar tools.
· Manage Kubernetes-based and containerized environments.
· Define and operate SLOs, SLIs, error budgets, dashboards, runbooks, and alerting standards.
· Implement observability, monitoring, and incident response practices.
· Participate in on-call rotations and respond to production incidents.
· Collaborate with engineering teams to improve automation, scalability, and platform resilience.
· Conduct postmortem reviews and drive continuous reliability improvements.
Required Skills & Qualifications
· Bachelor’s degree in Computer Science, Engineering, Information Systems, Software Engineering, or a related technical field, or equivalent practical experience.
· 6+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, DevOps Engineering, Backend Engineering, or Production Engineering.
· Strong software engineering skills in at least one language such as Python, Go, Java, TypeScript, or C#.
· Strong understanding of distributed systems, microservices, APIs, asynchronous processing, queues, databases, caching, retries, idempotency, and failure modes.
· Experience with cloud infrastructure on AWS, Azure, or GCP.
· Experience with Kubernetes, containers, Terraform or similar IaC tooling, CI/CD pipelines, and Linux-based systems.
· Experience with observability tools such as Datadog, Prometheus, Grafana, OpenTelemetry, CloudWatch, New Relic, Splunk, or Sentry.
· Experience defining and operating SLOs, SLIs, error budgets, alerting standards, dashboards, runbooks, and incident response practices.
· Strong communication skills and experience working across cross-functional teams.
Preferred Skills
· Cloud, Kubernetes, Infrastructure, Reliability Engineering, Security, or DevOps certifications.
· Experience in logistics, transportation, final-mile delivery, field-service software, routing, dispatch, or fleet operations.
· Experience working with operational SaaS or marketplace platforms.
· Experience driving automation, platform reliability, and operational excellence initiatives.
Posted On: 14-08-2026
At Coforge, we hire professionals based solely on their skills and qualifications and do not discriminate based on age, disability, religion, gender, sexual orientation, socioeconomic status, or nationality.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Site Reliability Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!