The Site Reliability Engineer will manage platform health through active monitoring, incident response, and the implementation of observability tools. They will also focus on reducing operational toil by building automation and maintaining reliable CI/CD pipelines.
The SRE Lead will define and implement the platform's reliability strategy, including observability, incident management, and NSOC operations. They will also lead and mentor a team of Site Reliability Engineers to ensure high availability and performance for a large-scale financial platform.