You will manage and stabilize the Middleware Platform, Azure, and Windows server estate by responding to alerts and resolving incidents within SLA. Additionally, you will drive automation efforts using scripting and AI agents to eliminate repetitive manual operational tasks.
Systems Engineer
Middleware Platform, Azure, Microsoft Windows Server, Monitoring, Automation and AI
Location: Remote — Pune or Baroda
Experience: 4–6 years
Shift: 24×7 rotation, including planned maintenance windows and change events
Role summary
You will keep our Middleware Platform, Azure and Windows server estate running day to day — responding to alerts, resolving incidents within SLA, Product/Application deployment, executing changes cleanly, and steadily improving the stability of the systems in your care.
You work within established processes and alongside Team Engineers and Database Administrators, taking ownership of your own queue and escalating appropriately.
This role also carries a real automation mandate. We expect you to notice repetitive manual work and remove it — through scripting, through Azure automation tooling, and increasingly through AI agents that handle routine operational tasks.
Experience
4–6 years in systems administration, Application and infrastructure support or cloud operations.
Minimum 4 years hands-on production Azure administration (not lab, training or certification-only exposure).
Minimum 3 years administering Windows Server in production.
Experience working to SLAs in a ticket-driven support environment.
Willing and available to work rotating 24×7 shifts and out-of-hours maintenance windows.
Preferred qualifications
Graduation or Above
Healthcare or other regulated-industry experience.
Key responsibilities
Azure platform administration
Administer Azure infrastructure day to day; provision, resize and decommission Azure resources following documented standards and approved change requests.
Support Azure PaaS workloads including App Services — deployments, configuration, scaling and restarts.
Apply resource tagging, cost-optimisation recommendations and policy compliance actions as directed.
Alert monitoring and response
Own the alert queue on your shift. Monitor Logic Monitor, Azure Monitor, Log Analytics and the enterprise monitoring platform; acknowledge, triage and action alerts within SLA.
Investigate alerts using KQL queries in Log Analytics, Application Insights telemetry and Windows Event Logs.
Identify duplicate or non-actionable alerts and propose tuning — threshold changes, aggregation windows, suppression rules. Flagging a bad alert is part of the job, not an optional extra.
Maintain monitoring coverage and visibility: report agents that are down or resources missing from monitoring, and build/maintain dashboards and workbooks that surface system health to the team.
System stabilisation
Resolve incidents end to end; support the Team Members during major incidents with diagnostics, evidence gathering and communication updates.
Contribute to root cause analysis — collect logs and timelines, document findings, and implement the corrective actions assigned to you.
Perform proactive health checks across applications, servers and Azure resources: disk, CPU, memory, service state, certificate expiry, backup success.
Build knowledge of application architecture, help create and update application and alert-handling documentation, and perform pre/post product functionality checks.
Track recurring issues in your queue and raise them as problem records rather than repeatedly applying the same workaround.
Perform performance troubleshooting using standard OS and Azure diagnostic tools.
AI agent development and process automation
Write and maintain PowerShell scripts for routine administration: user and resource provisioning, health checks, reporting, scheduled tasks and cleanup jobs.
Optional: Build automation using Azure Automation Runbooks, Logic Apps and Power Automate, with proper authentication via managed identity and alerting on failure.
Contribute to building AI agents that support operations — log summarisation, alert triage and classification, drafting incident notes, and answering questions from our runbook and knowledge base.
Apply prompt engineering to generate scripts, configurations and documentation, and verify AI output before it reaches production — you are accountable for what you deploy, regardless of what produced it.
Follow responsible-AI practice without exception.
Windows and application support
Administer Windows Server in Azure and hybrid environments
Support IIS — application pool management, bindings, virtual directories, logs and basic performance troubleshooting.
Perform SSL/TLS certificate tasks: renewal, installation, binding, store management and expiry monitoring.
Support high-availability configurations including Windows Failover Clustering and SQL Server Always On, working with the team engineers and the database administrators.
Process and collaboration
Work within ITIL incident, problem and change management using ServiceNow; raise and document change requests accurately and complete them within approved windows.
Maintain and improve runbooks, SOPs and knowledge articles — especially after resolving something the documentation didn't cover.
Provide clear, timely shift handovers and status updates to stakeholders during incidents.
Communication
Clear written and spoken English; able to write an incident update a non-technical stakeholder can follow.
Reliable, organised, and comfortable asking for help early rather than late.
Technical
Azure PaaS and IaaS administration, Windows Server administration, IIS administration and SSL certificate handling.
Azure Monitor and Log Analytics, with working ability to read and write basic KQL.
PowerShell scripting — able to write and modify scripts independently, not only run existing ones.
Ticketing and ITIL process discipline.
Practical use of AI/LLM tools in a work context, with at least one concrete example of manual work you've automated or removed.
Certifications
Required: AZ-104 (Azure Administrator Associate) — held, not "training completed".
Preferred: AZ-900, AI-900, ITIL Foundation, MCSA or equivalent Windows Server certification.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”