Design, scale, and maintain multi-cloud and on-premise infrastructure using Infrastructure as Code. Lead incident response, perform root cause analyses, and ensure system reliability and security.
\n
What you will do
Design, scale, and maintain Sysdig’s multi-cloud (AWS, GCP, Azure) and on-premise infrastructure using Infrastructure as Code.
Drive system reliability, performance, and scalability in close collaboration with software development teams.
Participate in on-call rotations, lead incident response, perform root cause analyses, and engineer long-term preventive solutions.
Enforce robust cloud security, data protection standards, and regulatory compliance across all environments.
What you will bring with you
5+ years of hands-on experience designing and managing large-scale, high-availability production infrastructure.
Strong expertise in container orchestration and runtime technologies, specifically Kubernetes and Docker.
Demonstrated track record of automating operational workflows to eliminate toil and boost system efficiency.
Proficiency with major cloud providers (AWS, GCP, IBM or Azure) and enterprise observability/monitoring frameworks.
What we look for
Advanced scripting or software development skills in Python, Go, or Bash within Linux environments.
Deep foundational knowledge of distributed systems, microservices architectures, and cloud-native network topology.
A proactive problem-solver dedicated to continuous improvement, automation, and operational excellence.
Adaptable self-starter who thrives, learns quickly, and delivers results in fast-paced, evolving environments.
When you join Sysdig, you can expect
Extra days off to prioritize your well-being
Mental health support for you and your family through the Modern Health app
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”