As a DevOps & Platform Operations Engineer, you will be responsible for the hands-on technical operations and reliability of the platform, ensuring it is securely deployed, monitored, backed up, and running smoothly in production. You’ll manage deployments, troubleshoot server and integration issues, oversee infrastructure upgrades, and respond to incidents by restoring services and identifying root causes.
The role is focused on keeping the platform secure, available, fast, and reliable while working closely with the product and development functions.
Company Profile:
Our client is an Australian-based claims management company focused on transforming the way insurance claims are managed and settled. Combining extensive industry expertise with technology and dedicated customer service, they deliver efficient, accurate, and transparent claims solutions for insurance companies and their customers.
They are currently seeking for an experienced DevOps & Platform Operations Engineer who is technically capable, detail-oriented, proactive, and energetic, with solid experience in DevOps, cloud infrastructure, platform operations, and maintaining reliable and scalable technology environments.
This is an excellent opportunity to build a long-term career with a growing company, working in a collaborative environment with opportunities to contribute to technology improvements and strengthen platform reliability.
Duties and Responsibilities:
Cloud & Server Management
- Manage the SaaS platform’s production and staging infrastructure, including cloud servers and services, databases, domains and DNS, SSL certificates, storage, networking, environment configuration, application services, user access and permissions, and infrastructure scaling. Monitor system health and ensure infrastructure remains appropriately configured as the SaaS platform grows
Deployment & Release Management
- Take approved the SaaS platform’s code and reliably deploy it into staging and production environments
- Develop and maintain automated CI/CD deployment pipelines so releases become increasingly automated and repeatable. Maintain clear separation between Development → Staging → Production. Ensure releases can be rolled back quickly if problems occur
Monitoring & Uptime
- Implement monitoring and alerting across the SaaS platform’s environment — application availability, server health, database performance, API performance, error rates, storage, CPU/memory utilization, failed processes, integration failures, security events. Where possible, identify problems before users report them
Incident Response
- Act as the first technical point of contact when the SaaS platform experiences an operational issue (website unavailable, application errors, server failures, database problems, failed deployments, API failures, authentication issues, email/notification failures, performance degradation, third- party integration problems). Diagnose the issue, restore service and document the root cause
- For significant incidents, produce a short Root Cause Analysis explaining: what happened → why it happened → how it was fixed → how recurrence will be prevented
Backup & Disaster Recovery
- Own the SaaS platform’s backup and recovery systems. Ensure databases and critical files are automatically backed up, backups are geographically/reliably stored, recovery procedures are documented, backups are periodically tested, and production environments can be rebuilt if necessary. Maintain a documented Disaster Recovery Procedure
Security
- Maintain good cloud and application security practices — access control, multi-factor authentication, secrets management, API credentials, encryption, firewall/security configuration, dependency and vulnerability monitoring, security updates, logging, production access management, backup security. Because the platform handles insurance claim and customer information, security and data protection must be treated as core operational requirements
Requirements
- At least 5 years of hands-on experience in DevOps, Platform Operations, Cloud Engineering, or a similar role
- Strong experience with Microsoft Azure, including infrastructure, monitoring, networking, security, and production deployments
- Experience with Microsoft Entra ID, access controls, and identity management
- Experience managing Linux servers, web applications, and production environments
- Experience with Git/GitHub, CI/CD pipelines, Docker, and application deployments
- Knowledge of DNS, SSL/TLS, SQL databases, backups, REST APIs, and third-party integrations
- Experience with cloud monitoring, logging, incident response, troubleshooting, and root-cause analysis
- Understanding of infrastructure and application security
- Experience with Python, Bash, PowerShell, or similar scripting languages
- Strong problem-solving skills with a methodical and hands-on approach
- Comfortable working independently, taking ownership, and communicating technical issues in plain English
- Comfortable using AI tools as part of everyday work
- Demonstrates enthusiasm, creativity, and genuine passion for the role and the work
- Shows a high level of engagement, initiative, and ownership in completing tasks and contributing to team objectives
- Proactive in identifying opportunities, solving problems, and contributing ideas beyond assigned responsibilities
- Brings a positive, collaborative, and results-oriented mindset to the workplace
Advantageous or Nice-to-Have Skills/Experience:
- Terraform or other Infrastructure as Code experience