The Site Reliability Engineer will manage platform health through active monitoring, incident response, and the implementation of observability tools. They will also focus on reducing operational toil by building automation and maintaining reliable CI/CD pipelines.
The Role
The Site Reliability Engineer is the hands-on operator and automation builder who keeps the platform running. Reporting to the SRE Lead, you will be responsible for the day-to-day health of the platform --- monitoring, alerting, incident response, and the continuous automation work that makes the platform easier to operate and harder to break. You will work closely with the DevSecOps engineers, cloud engineers, and security team to ensure that reliability is built into every layer of the platform, not bolted on afterwards. This is a demanding role on a high-stakes platform --- and a rare opportunity to do SRE work that genuinely matters at scale.
What You'll Do
NSOC Operations & Monitoring --- Work as part of the 24/7 Network & Security Operations Centre (NSOC) --- actively monitoring the platform's health across all services, cloud environments, and network layers. Triage incoming alerts, distinguish signal from noise, escalate genuine incidents promptly, and maintain the situational awareness that keeps the team ahead of problems before they become outages.
Incident Response & Triage --- Serve as a first responder for platform incidents --- detecting, triaging, and containing issues in real time. Execute runbooks under pressure, coordinate with engineering teams during active incidents, and own your part of the resolution. After every significant incident, contribute to the blameless post-mortem process --- documenting what happened, what worked, and what needs to change.
Observability Implementation & Maintenance --- Implement, maintain, and continuously improve the platform's observability stack --- including centralised logging, metrics collection (Prometheus, Datadog, or equivalent), distributed tracing (Jaeger, OpenTelemetry, or equivalent), and alerting rules. Instrument new services from day one, tune alerting thresholds to reduce noise without missing real issues, and build dashboards that give the team the operational visibility they need.
Automation & Toil Reduction --- Identify repetitive operational tasks and build the automation that eliminates them. Write scripts, tools, and workflows in Bash, Python, or equivalent to reduce toil --- replacing manual processes with reliable, repeatable automation. Toil that exists today should not exist next quarter. This is an ongoing discipline, not a project with an end date.
Reliability & Chaos Testing --- Support the SRE Lead in running the platform's reliability engineering programme --- executing load tests, stress tests, and chaos engineering experiments to validate the platform's resilience under real conditions. Document findings, flag vulnerabilities, and work with engineering teams to address weaknesses before they become production incidents.
Runbook Development & Maintenance --- Write, test, and maintain operational runbooks for all known failure modes and recovery procedures. A runbook is only valuable if it works --- regularly test each runbook against real or simulated conditions, update it when the platform changes, and make sure every procedure is accurate and actionable. Runbooks are not documentation --- they are operational tools that must be ready to use at 2am.
SLO Monitoring & Reporting --- Monitor SLO attainment and error budget consumption across all platform services on an ongoing basis. Proactively flag services approaching error budget limits to the SRE Lead and relevant engineering teams. Track and report DORA metrics --- deployment frequency, lead time, change failure rate, and MTTR --- as part of the team's ongoing performance visibility.
Platform Health & Capacity Monitoring --- Continuously monitor resource utilisation, service performance, and infrastructure health across both cloud environments (AWS and Azure). Identify trends that suggest emerging capacity constraints or degrading performance, and surface these findings to the SRE Lead and Platform Manager before they become incidents.
CI/CD Pipeline Health --- Monitor the health and performance of the platform's CI/CD pipelines in collaboration with the DevSecOps team. Identify pipeline failures, flaky tests, or degraded build performance and escalate promptly. Reliable pipelines are part of platform reliability --- treat them accordingly.
Documentation & Knowledge Sharing --- Maintain accurate, up-to-date operational documentation --- procedures, architecture notes, incident timelines, and lessons learned. Share knowledge actively with your SRE colleagues and across the broader engineering team. A well-documented platform is a more reliable platform.
What We're Looking For
Must Have
2+ years of SRE, DevOps, or platform operations experience in a production environment --- with hands-on responsibility for monitoring, alerting, and incident response.
Hands-on experience with observability tooling --- logging (ELK, Loki, Datadog, or equivalent), metrics (Prometheus, Datadog, or equivalent), and alerting. You have built dashboards and tuned alert rules, not just viewed them.
Experience responding to production incidents --- triaging, containing, and resolving issues under pressure. You understand incident severity classification and escalation paths.
Scripting skills in Bash and Python (or equivalent) --- for writing automation, operational tooling, and toil-reduction scripts.
Solid understanding of cloud infrastructure on AWS and/or Azure --- compute, managed services, networking, and the platform-level health indicators you should be monitoring.
Genuine care about reliability --- you understand why SLOs and error budgets matter, and you approach operational work with the discipline that high-availability financial services demand.
Nice to Have
Experience with distributed tracing --- Jaeger, OpenTelemetry, AWS X-Ray, or equivalent --- and using trace data to diagnose performance and reliability issues.
Exposure to chaos engineering tools --- Chaos Monkey, Gremlin, or equivalent --- or experience running load and stress tests in production or pre-production environments.
Familiarity with Kubernetes operational patterns --- pod health, cluster observability, resource constraints, and Kubernetes-native alerting and scaling behaviours.
Experience in a regulated environment --- fintech, banking, or similar --- with an understanding of the uptime, audit, and compliance expectations financial services carry.
AWS and/or Microsoft Azure certification --- SysOps, Cloud Practitioner, or equivalent entry-to-mid-level cloud certifications.
Familiarity with DORA metrics and what they tell you about engineering team health and deployment risk.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”