The SRE Lead will define and implement the platform's reliability strategy, including observability, incident management, and NSOC operations. They will also lead and mentor a team of Site Reliability Engineers to ensure high availability and performance for a large-scale financial platform.
About the Engagement
We are building a next-generation digital banking platform for one of Africa's largest financial institutions --- a product built to serve 15 million users at launch, scaling to hundreds of millions across the continent and diaspora. On a financial platform at this scale, reliability is not a feature --- it is a promise. This is a once-in-a-generation opportunity to build the reliability function that keeps that promise.
The Role
The SRE Lead is the owner of the platform's reliability function --- the person who defines what reliability means on this platform, builds the systems and processes to deliver it, and leads the team that keeps it. Reporting to the Platform Manager and leading a team of Site Reliability Engineers, you will design and operate the platform's observability stack, incident management framework, and NSOC --- ensuring the platform is always visible, always monitored, and always recoverable. On a banking platform serving millions of Africans, downtime is not just a technical failure --- it is a failure of trust. You are the person who makes sure that never happens.
What You'll Do
SRE Strategy & Reliability Standards --- Own the platform's Site Reliability Engineering function --- defining the reliability philosophy, standards, and practices that govern how the platform is operated. Establish and enforce SLOs and error budgets for all platform services, working with engineering teams to ensure reliability targets are defined, measured, and taken seriously. Build a culture where reliability is a shared engineering responsibility, not just an ops concern.
NSOC Design & Operations --- Design, build, and operate the platform's 24/7 Network & Security Operations Centre (NSOC) --- the central nervous system of the platform's operational awareness. Define the monitoring coverage, alert routing, escalation paths, and on-call workflows that ensure no incident goes undetected and no alert goes unacted upon. The NSOC must be staffed, tooled, and operating continuously from the moment the platform goes live.
Observability Architecture & Implementation --- Own the platform's observability stack --- designing and implementing the centralised logging, metrics collection, distributed tracing, and alerting infrastructure that provides complete visibility into every layer of the platform. Ensure every service is instrumented from day one in production, alerting is tuned to minimise noise without missing signal, and dashboards give the team the operational intelligence they need at a glance.
Incident Management --- Own the platform's end-to-end incident management process --- from detection and triage through containment, resolution, and post-incident review. Define the on-call framework, escalation paths, and severity classification standards. Ensure every significant incident results in a blameless post-mortem with clear action items tracked to completion. Continuously improve incident response speed and quality --- tracking MTTR as a primary platform health metric.
Reliability Engineering & Chaos Testing --- Design and run a proactive reliability engineering programme --- including chaos engineering experiments, load and stress testing, failure mode analysis, and disaster recovery drills. Do not wait for production to reveal failure modes. Identify and address weaknesses in the platform's resilience before they become incidents, and validate that the platform's recovery capabilities work exactly as designed under real conditions.
SLO & DORA Metrics --- Own the collection, baselining, and reporting of the platform's reliability and engineering performance metrics --- including SLO attainment, error budget consumption, deployment frequency, lead time for changes, change failure rate, and mean time to recovery (MTTR). Surface these metrics to the Platform Manager in real time and use them to identify where engineering practice and platform reliability need focused improvement.
Toil Reduction & Automation --- Systematically identify, measure, and eliminate toil --- the repetitive, manual, automatable operational work that consumes engineering capacity without improving reliability. Build automation that makes the platform easier to operate at scale and frees the SRE team to focus on reliability engineering rather than routine operations. Toil reduction is an ongoing discipline, not a one-time project.
Runbook Development & Knowledge Management --- Own the platform's operational runbooks --- ensuring every known failure mode, recovery procedure, and escalation path is documented, version-controlled, regularly tested, and immediately accessible during an incident. The runbook library must be current, complete, and something an engineer can actually use under pressure at 2am.
Capacity Planning --- Work with the Platform Manager and Cloud Engineers to continuously monitor platform resource utilisation and anticipate capacity requirements before they become constraints. Ensure the platform can absorb growth --- in users, traffic, and data volume --- ahead of demand, not in response to it.
Team Leadership --- Lead, mentor, and develop the Site Reliability Engineers on the team --- setting clear expectations, reviewing their work, and creating the conditions for them to grow into increasingly independent, senior practitioners. Build a high-performing SRE team that the rest of the engineering organisation trusts and relies on.
What We're Looking For
Must Have
5+ years of SRE, DevOps, or platform engineering experience --- with at least 2 years in a lead or senior individual contributor role owning a reliability function in production.
Deep hands-on experience with observability tooling --- centralised logging, metrics (Prometheus, Datadog, or equivalent), distributed tracing (Jaeger, Open Telemetry, or equivalent), and alerting frameworks. You have built and operated an observability stack, not just used one.
Proven experience designing and running incident management processes --- including on-call frameworks, severity classification, escalation paths, blameless post-mortems, and action item tracking.
Strong understanding of SRE principles --- SLOs, SLAs, error budgets, toil reduction, and reliability engineering as a discipline distinct from DevOps.
Hands-on experience with cloud infrastructure on AWS and/or Azure --- including compute, managed services, networking, and the observability capabilities of both platforms.
Experience with DORA metrics --- deployment frequency, lead time, change failure rate, MTTR --- and using engineering performance data to drive improvement.
Strong scripting and automation skills --- Bash, Python, or equivalent --- for building the tooling and automation that makes the platform easier to operate at scale.
Nice to Have
Experience designing and running chaos engineering programmes --- Chaos Monkey, Gremlin, or equivalent --- and using failure injection to validate platform resilience.
Experience building and operating a 24/7 NSOC or SOC function in a regulated financial services environment.
Familiarity with container and Kubernetes operational patterns --- pod health, cluster observability, resource quotas, and Kubernetes-native alerting.
Experience with capacity planning and FinOps --- using utilisation data to inform infrastructure scaling and cost decisions.
AWS and/or Microsoft Azure certification --- SysOps, DevOps, or Solutions Architect tracks.
Experience working in fintech, banking, or a regulated environment --- with an understanding of the reliability and availability expectations that financial services demand.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”