For Employers
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

The Platform Engineer will build and maintain monitoring, observability, and infrastructure automation to ensure cloud platform reliability. Responsibilities include managing alert routing, deploying infrastructure-as-code, and acting as a first responder for system incidents.

As a Platform Engineer at ThinkOn, you'll build and maintain the monitoring, observability, and infrastructure automation that keeps our cloud platform reliable. Your day-to-day will span Zabbix and Prometheus/Grafana for monitoring and dashboards, Opsgenie for alert routing and on-call management, and infrastructure-as-code tooling (Ansible, Terraform, GitLab CI/CD) to deploy and manage it all. You'll work within a VMware Cloud Foundation (VCF) and Kubernetes environment, collaborating with infrastructure, network, security, and DevOps teams.

ThinkOn is a remote-first organization. At this time, we are welcoming candidates from Canada for this position.

Please note that this listing is for a current vacancy at ThinkOn. We are looking for qualified candidates who are eager to contribute and grow with us. If you're interested in this opportunity, we encourage you to apply soon, as we are reviewing applications on a rolling basis.

You Will:

Monitoring & Observability

  • Deploy, configure, and maintain Zabbix for system and network monitoring across the platform.
  • Build and maintain Prometheus exporters and Grafana dashboards for capacity planning, performance metrics, and operational visibility.
  • Configure and manage Ops genie for alert routing, escalation policies, and on-call schedules.
  • Analyze alert noise, tune thresholds, and reduce false positives to keep alerting actionable.
  • Integrate monitoring systems with ticketing, communication, and incident management tools.
  • Maintain and optimize monitoring infrastructure — database tuning, storage management, high availability.

Infrastructure Automation & CI/CD

  • Write and maintain Ansible playbooks for deploying and configuring monitoring infrastructure.
  • Use Terraform for provisioning infrastructure resources where applicable.
  • Build and maintain GitLab CI/CD pipelines for automated testing, linting, and deployment of monitoring and infrastructure code.
  • Follow infrastructure-as-code practices — version-controlled, peer-reviewed, reproducible.

Incident Response & Support

  • Act as a first responder for monitoring-related incidents and alerts.
  • Investigate and resolve performance issues, outages, or anomalies detected by monitoring systems.
  • Escalate to the appropriate teams (network, security, infrastructure) when needed.
  • Document incidents, root causes, and resolutions. Contribute to post-incident reviews.
  • Provide technical support to internal teams on monitoring tools and dashboards.

Platform Operations

  • Work within VMware vCenter / VCF and Kubernetes environments to support monitoring and infrastructure needs.
  • Manage notification infrastructure (SMTP relay configuration, delivery troubleshooting).
  • Support compliance requirements (ISO 27001, SOC 2) by maintaining audit logging, access controls, and security configurations for monitoring systems.

You Have:

  • Diploma or degree in Computer Science, IT, or a related field (or equivalent practical experience).
  • Eligible to obtain Secret Level Clearance within your first 3 months of employment. This requires the successful candidate to be a Canadian Citizen and have 10+ years of verifiable police background history.
  • Relevant certifications are a plus but not required (e.g., Zabbix Certified Professional, CKA, CompTIA Linux+, ITIL v4 Foundations).
  • Hands-on experience with Zabbix(or comparable: Nagios, Icinga,Checkmk).
  • Working knowledge of Prometheus and Grafana— writing exporters, building dashboards, PromQL.
  • Experience with alert management and on-call tooling(Ops genie, PagerDuty, or similar).
  • Comfort with Linux systems administration (this is a Linux-heavy environment).
  • Proficiency in scripting and automation— Bash and Python at minimum.
  • Experience with at least one IaCtool (Ansible, Terraform).
  • Familiarity with CI/CD pipelines(GitLab CI, GitHub Actions, Jenkins, or similar).
  • Understanding networking fundamentals— TCP/IP, DNS, SNMP, bandwidth/latency concepts.
  • Basic database administration (PostgreSQL or MySQL) for monitoring tool backends.
  • Strong diagnostic and troubleshooting skills — you can work through a problem methodically.
  • Clear written and verbal communication —you'll document your work and explain technical issues to varied audiences.
  • Attention to detail — monitoring generates a lot of data, and you need to separate signal from noise.
  • Comfort working independently in a remote environment while collaborating across teams.
  • Ability to stay composed during incidents and work under time pressure.

Nice-to-Have

  • Experience with Kubernetes operations and troubleshooting.
  • Familiarity with VMware vSphere / VCF environments.
  • Exposure to log aggregation tools (ELK/OpenSearch, Loki, Graylog).
  • Knowledge of email security standards (SPF, DKIM, DMARC).
  • Experience cloud monitoring in multi-tenant or service-provider environments.
  • Understanding of Canadian data sovereignty requirements or public-sector compliance (PIPEDA, ITSG-33/PBMM).
  • Familiarity with ITSM platforms like Service Desk Plus by Managed Engine, ZenDesk, ServiceNow, Jira Service Management, or similar.

Benefits and Perks: 

  • A remote-first culture 
  • Competitive compensation package
  • Flexible time-off for vacation, plus illness & personal days
  • Comprehensive health and dental benefits, including GRSP & 401k Matching Program

About ThinkOn Inc. | Where Data Thrives:

Founded in 2013, ThinkOn is a managed infrastructure services provider (MISP) with a global data center footprint, focused on empowering partners to do whatever they need to do with their data. ThinkOn’s team of data-obsessed experts protect clients’ data like it’s their own, making it more resilient, secure, actionable, and searchable. ThinkOn is Channel First and works with a global network of value-add resellers and managed service providers to provide creative, turnkey Infrastructure-as-a-Service (IaaS), Disaster Recovery-as-a-Service (DRaaS), and Backup-as-a-Service (BaaS) solutions and data management services that are fast, flexible, scalable, highly secure, and cost-effective with predictable pricing and no hidden fees. ThinkOn has data centers located across North America, the United Kingdom, Australia, and the Caribbean.

Recognized for its substantial growth and global success, ThinkOn has been named to several notable lists, including The Globe and Mail’s Report on Business ranking of “Canada’s Top Growing Companies” (2021 and 2022), the Deloitte Technology Fast 500 (2021 and 2022), the Deloitte Technology Fast 50 (2021), CIOReview’s “Most Promising Backup Solution Provider” (2021), the Canadian Business “Growth List,” and Channel Daily News’ “Top 100 Solution Providers” (2021).

ThinkOn is headquartered in Toronto, Ontario. To learn more, visit www.ThinkOn.com

Accessibility Accommodations:

ThinkOn is committed to a workforce that is reflective of diverse populations. We welcome applications from qualified individuals from all backgrounds. In accordance with the Accessibility for Ontarians with Disabilities Act (AODA) and accessibility standards across Canada, ThinkOn provides accommodations to job applicants with disabilities throughout the recruitment process. If you require accommodations, please let us know and we will work with you to meet your needs. We are committed to a selection process and work environment that is inclusive, equitable, accessible, and adheres to our corporate values.

Learn more about our corporate values at www.ThinkOn.com/values/.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Senior Cloud / DevOps Engineer - USA

Full Time United States $115K - $130K per year Software Development

Principal Data Analyst

Full Time Brazil Software Development

Senior Software Engineer - FinOps

Full Time United States $104K - $174K per year Software Development

Engineering Lead (m/w/d) TS, React, Next.js & AI 100% Remote

Full Time Germany €90000 - €120K per year Software Development

Junior Fullstack Engineer (m/w/d) TS, React, Next.js & AI 100% Remote

Full Time Germany €50000 - €65000 per year Software Development

Staff Software Engineer, Communication & Connectivity

Full Time United States $204K - $255K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Platform Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified