Senior Site Reliability Engineer (Performance and Scalability)

 Posted 2 hours ago
  
 Egypt
Apply Now

Please mention DailyRemote when applying

Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself.

What you'll do

  • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load.
  • Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results.
  • Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale.
  • Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.
  • Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on.
  • Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems.

Requirements

What you'll bring

  • 5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minute.
  • A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted.
  • Deep AWS experience and a solid grasp of Postgres performance and scaling.
  • Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar.
  • A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep.

Benefits

  • Immediate, large-scale impact on a high-growth business
  • Top-of-the-market compensation packages
  • Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more

Similar Jobs

See all Remote Others jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Site Reliability Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified