Please mention DailyRemote when applying
Location : Work from Home - Province of Quebec
The Site Reliability Specialist on the IT Operations team contributes to the reliability, availability, performance, and resilience of Sherweb's platforms and services.
This is a highly technical individual contributor role that applies Site Reliability Engineering (SRE) principles to production environments. The role combines systems administration, software development, automation, observability, and operational excellence to improve service reliability, reduce operational toil, and increase platform scalability.
Working closely with Infrastructure, Development, DevOps, Platform, Security, and Product teams, the Site Reliability Specialist helps ensure production systems remain stable, supportable, and continuously improving through engineering and automation practices.
Here's how you will contribute to the success of the company
· Develop, maintain, and improve scripts, automation workflows, and operational tooling to reduce manual intervention, improve reliability, and lower operational toil.
· Apply SRE principles to improve the reliability, availability, performance, and resilience of Sherweb’s hosted platforms and production services.
· Implement and support reliability standards, service level objectives (SLOs), service level indicators (SLIs), and operational practices established for platforms and services.
· Use a developer mindset to transform repetitive operational tasks into scalable, reusable, documented, and supportable automation.
· Provide advanced operational support and resolve incidents affecting production services while collaborating closely with SRE, Infrastructure, Development, DevOps, and Product teams.
· Investigate recurring issues and perform root cause analysis to identify short-term corrective actions and long-term reliability improvements.
· Build, support, and maintain production systems and hosted service technologies while following operational procedures, security best practices, and compliance requirements.
· Improve monitoring, alerting, logging, metrics, and operational visibility to help the team detect issues earlier, understand system behavior, and prevent incidents.
· Contribute to improving end-to-end observability and system understanding through metrics, logs, traces, telemetry, and operational diagnostics.
· Contribute to observability-as-code, infrastructure-as-code, configuration-as-code, and automation practices where applicable.
· Explore and leverage Azure AI Foundry, Power Automate, and AI agent capabilities to improve operational efficiency, automate repetitive workflows, and accelerate incident response or service reliability improvements.
· Participate in platform lifecycle activities, deployments, maintenance windows, migrations, and continuous service improvement initiatives.
· Collaborate with developers, architects, subject matter experts, DevOps, and infrastructure teams to support implementation, optimization, troubleshooting, and operational readiness of services.
· Create and maintain operational documentation, including SOPs, runbooks, troubleshooting guides, automation documentation, maintenance procedures, and knowledge-sharing materials.
· Track, organize, and manage incidents, requests, and service tickets to respect SLAs and ensure clear communication through the ITSM process.
· Participate in rotational on-call duty and perform maintenance work outside normal business hours when required.
· Carry out all other related tasks per the job’s evolution and departmental needs.
Here's what you need to have and master to get the job
Education
· College or university degree in computer science, software development, information technology, engineering, or a combination of equivalent training and experience.
Experience
· 3 to 5 years of experience in systems administration, IT operations, infrastructure support, software development, DevOps, automation, or a similar technical role.
· Experience supporting production systems in business-critical and customer-facing environments.
· Proven experience improving operational efficiency through automation and engineering practices.
Core Skills
· Strong scripting or development skills with at least one language such as PowerShell, Python, Bash, JavaScript, TypeScript, or C#.
· Proven ability to design, write, test, troubleshoot, document, and maintain scripts or small applications used to automate operational tasks.
· Proven experience supporting Microsoft and/or Linux server environments, including troubleshooting, maintenance, operational support, and automation.
· Strong diagnostic, investigation, and problem-solving skills with the ability to analyze incidents, identify root causes, and implement sustainable improvements through automation or engineering practices.
· Good understanding of distributed systems, networking concepts, system dependencies, availability, performance, reliability, and service operations in production environments.
· Experience with monitoring, alerting, observability, log management, telemetry and operational data analysis to detect, troubleshoot, and prevent issues.
· Familiarity with version control, Git-based workflows, CI/CD pipelines, code review practices, and deployment automation.
· Experience with infrastructure as code, configuration management, or automation tools such as Terraform, Ansible, DSC, Azure DevOps, GitHub Actions, Docker, or Kubernetes is an asset.
· Familiarity with Azure AI Foundry, Power Automate, Copilot/AI agents, or agent-based automation concepts is an asset.
· Knowledge of high availability environments, virtualization, cloud services, backup and restore practices, and production support models is an asset.
Professional Attributes
· Autonomous, reliable, and motivated, with a continuous learning mindset and a strong interest in improving reliability through software engineering and automation practices.
· Strong communication and collaboration skills, with the ability to work effectively with technical and non-technical stakeholders.
· Excellent English skills, both spoken and written, are essential; fluency in French is an asset.
· Relevant industry certifications such as Microsoft Azure, Red Hat, Linux Foundation, Kubernetes, DevOps, or observability platforms are considered assets.
Additional Requirements
*Availability for rotation on the on-call schedule in a 24/7 environment.
Benefits of working at Sherweb
Sherweb is first and foremost a culture where our customers’ needs are at the heart of everything we do, supported by Sherwebers committed to living our values of passion, teamwork, and integrity.
Dynamic and flexible work environment
Flexible total compensation package
Unlimited growth opportunities
A vibrant social life (Sherweblife)
A rich calendar of virtual and in-person activities designed to foster connection throughout the year.
At Sherweb, we believe in transparency and pay equity. The salary range provided is intended to give an indication of what you can expect for this role. However, we recognize that each candidate brings a unique set of skills and experiences. The final compensation package will be tailored to reflect the selected candidate’s qualifications and expertise, ensuring we remain competitive and fair in our offers.
Reasons for the requirement of English: Sherweb has international customers and fluency in English is the only way to ensure proper service delivery to them. The main tasks related to this position require written and oral communication with an English-speaking clientele at all times.
#LI-Remote
#LI-SG1
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Development
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!