For Employers

Digital Zone

Senior Site Reliability Engineer (Performance and Scalability)

Posted 8 hours ago
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review

Questions interviewers often ask for this role, with sample answers.

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable

Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself.

What you'll do

  • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load.
  • Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results.
  • Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale.
  • Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.
  • Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on.
  • Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems.

Requirements

What you'll bring

  • 5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minute.
  • A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted.
  • Deep AWS experience and a solid grasp of Postgres performance and scaling.
  • Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar.
  • A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep.

Benefits

  • Immediate, large-scale impact on a high-growth business
  • Top-of-the-market compensation packages
  • Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Middle Business Analyst

Full Time Software Development

AI Internship

Internship Software Development

Application Architect - Remote from EU

Freelance Software Development

Software Engineer II - (Ruby)

Full Time Software Development

Senior Support Engineer II

Freelance Software Development

Junior Frontend Engineer

Full Time Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 215,359+ Jobs in Site Reliability Engineer

Answer easy questions

Answer easy questions

215,359+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified