Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
Perform manual Tier 1 triage by monitoring, detecting, and classifying incidents across a diverse application portfolio. Coordinate with support groups and application owners to resolve incidents while contributing to the shift toward proactive issue prevention.
About the Role
As part of Bell's AIOps program, AIOps Support Engineers
perform hands-on, manual triage of alarms and monitoring alerts across a large
and diverse application portfolio. The role starts with 50 applications and
grows to roughly 320 over two years, spanning different technology stacks,
different log aggregation tools, and different lifecycle dispositions (Invest,
Tolerate, Retire, Migrate). The focus today is disciplined, high-quality manual
triage with a clear roadmap toward more automated, proactive detection as the
underlying telemetry backbone (built on Open Telemetry) matures. This is not an
Agentic-automation role; it is the foundational human layer that makes
accurate, standardized data and eventually automation possible.
What You'll Do
• Perform manual Tier 1 triage: monitor,
detect, classify, and route incidents based on alarm and alert signals from a
range of source applications.
• Work across a diverse tool set: interpret alerts
from multiple log aggregation and observability platforms Dynatrace, New Relic,
ManageEngine, Glass box, and others and translate them into consistent,
actionable incident data.
• Support standardized, real-time telemetry help
build and maintain a clean Open Telemetry-based data backbone, ensuring alarm
and alert data is standardized, timely, and reliable.
• Move from reactive to proactive: contribute
to the shift from reactive incident response toward proactive issue prevention by
flagging patterns, gaps, and data-quality issues in monitoring coverage.
• Understand a complex app ecosystem: learn
and track the technology stack, architecture disposition (Invest, Tolerate,
Retire, Migrate), and ownership of each supported application.
• Engage support groups and app owners: coordinate
with the correct support group and application owner for each application to
resolve or escalate incidents quickly.
• Support a growing, moving target: adapt as
new applications are onboarded the portfolio will grow from 50 to approximately
320 applications over two years.
• Practice data stewardship: flag
inconsistent, noisy, or low-quality alert data and help drive it toward a
clean, standardized state.
Key Technical Competencies
• Application
& Production Support (L1/L1.5)
• Incident, Problem & Change Management (ITIL)
• Application, Middleware & Database Log
Analysis
• Monitoring & Observability (Dynatrace, New
Relic, AWS CloudWatch, and similar tools)
• OpenTelemetry & Distributed Tracing (working
knowledge)
• Proactive Alert Monitoring & Outage
Prevention
• Linux & Windows OS Troubleshooting
• Basic Networking (TCP/IP, DNS, HTTP/HTTPS, SSL,
Load Balancers)
• SQL & Database Query Analysis
• API & Integration Troubleshooting
• Hybrid
Environment Support (On-Prémises + Cloud)
• ServiceNow / Jira Ticket Management
What We're Looking For
• Experience: 2–5 years in application or
production support (L1/L1.5), preferably supporting hybrid on-premises and
cloud environments.
• Willingness to do manual, hands-on triage: comfortable
with the detail work of alert monitoring and classification, especially in the
early stages of the program before automation scales.
• Adaptability: able to work across a
complex, changing application ecosystems different technology stacks, different
log aggregation tools, and different architecture placements.
• Customer-centric mindset: focused on
service stability, operational excellence, and continuous improvement.
• Clear, fast communication: able to interpret
alert data with speed and clarity and turn it into actionable next steps for
support groups and application owners.
• Growth orientation: interested in
developing toward more proactive, automation-supported monitoring as the
program matures.
Must-Have Skills
• Understanding of the Incident Management
Lifecycle, and Problem, Change Request, and Service Request concepts (ITIL)
• Basic CMDB concepts
• Hands-on experience working with log aggregation
technologies (Dynatrace, New Relic, ManageEngine, Glassbox, or similar)
• Working knowledge of JSON and XML, and basic
file/task automation
• Knowledge of IT infrastructure and basic
networking: VMs, firewalls, load balancers, containers, OpenShift (OCP),
Kubernetes
• Unix Shell scripting and Windows batch file
creation
• Basic knowledge of cloud concepts, storage, and
security
• Basic understanding of TLS, SSL, tokens, and
secret management
• Familiarity with API gateways and API toolkits
(Postman, SOAP UI, or similar)
• Ability to support outage response and work
effectively with cross-functional teams
• Ability to contribute to Root Cause Analyses
(RCAs) and knowledge/runbook creation
• Basic understanding of SLA/SLO concepts,
reporting, and availability calculation
• Basic understanding of data concepts: data
latency, data fragmentation, data lineage, and data marts
Nice-to-Have Skills
• Familiarity with AI concepts such as prompt
engineering, knowledge graphs, and Retrieval-Augmented Generation (RAG)
• Exposure to public cloud platforms (AWS, Azure,
or GCP)
• Scripting or automation experience beyond basic
file automation (e.g., Python)
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Support Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”