Please mention DailyRemote when applying
See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewQuestions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
The engineer will own the health, reliability, and lifecycle management of the Azure virtual machine estate. They are responsible for implementing automation, managing backups and disaster recovery, and driving root cause analysis for infrastructure incidents.
We are seeking a highly experienced (8-10 years) Cloud Infrastructure Engineer to own the reliability of our Azure virtual machine estate. Predominantly Windows today, with a smaller Linux footprint, currently around 60 VMs and growing to 300+ as we bring more business units onto the platform. Your job is to keep it patched, backed up, monitored, and recoverable, and automate yourself out of the repetitive tasks. We expect problems to be solved once, in code, with Terraform and scripting.
· Own the health and reliability of the Azure VM estate (Windows and Linux), availability, performance, and capacity.
· Own / Run patching and update management with Azure Update Manager: patch compliance, maintenance windows, remediation of failures, and handling applications that need a version held or pinned without falling out of the compliance cycle.
· Own backup and disaster recovery: Azure Backup, Azure Site Recovery, and regular restore testing. A backup that was never restored doesn't count.
· Build monitoring and alerting with Azure Monitor and KQL (Kusto Query Language) and drive auto-remediation so known issues fix themselves.
· Automate VM lifecycle operations (provisioning, configuration, decommissioning) with Terraform and scripting.
· Manage incidents affecting the VM estate, drive root cause analysis, and fix the class of problem, not just the instance.
· Diagnose cases where the real root cause is security tooling. Antivirus/EDR flagging and quarantining an application file, for example, rather than assuming the fault sits in the application or the infrastructure.
· Support security-led investigations on the VM estate: pull logs, process activity, and access history on request, and hold off on remediating or restarting a box until security has cleared it.
· Reduce toil: identify repetitive manual work and eliminate it through automation.
· Document runbooks and operational standards, so the platform is operable by the whole team.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 215,794+ Jobs in Cloud Infrastructure Engineer
Answer easy questions
215,794+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”