Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
You will lead a team of onshore and offshore engineers to manage and improve supply chain systems, ensuring 24x7 reliability and performance. Responsibilities include setting Service Level Objectives, driving incident response, and fostering a culture of automation and continuous development.
With a career at The Home Depot, you can be yourself and also be part of something bigger.
Position Purpose:
As Manager, Reliability Engineering for Transportation and Delivery Fulfillment, you lead the team that runs and improves the supply chain systems that move freight to our stores and distribution centers and get product to our customers. Your team spans onshore and offshore engineers, including staff engineers, and you hold operational and engineering work to a single set of goals. You own these systems 24x7. When a major incident is declared, you are the escalation point: you open the call, direct the technical response, and tell business and executive stakeholders about the impact, the scope, and the expected recovery. Those updates go out on a cadence you publish in advance. You run both domains as one practice, with one on-call and escalation model, one problem review, and one prioritized reliability backlog. You map the Critical User Journeys (CUJs) the business depends on, then set and enforce Service Level Objectives (SLOs) for availability and performance against them. Problem management, automation, and applied AI are how you remove recurring operational work for good. You keep changes from causing incidents, and you deliver resilience testing and security remediation on the dates we commit to. You hire, develop, and recognize your engineers.
Key Responsibilities:
30% Delivery & Execution:
10% Support & Enablement:
50% People:
10% Learning:
Direct Manager/Direct Reports:
Travel Requirements:
Physical Requirements:
Working Conditions:
Minimum Qualifications:
Preferred Qualifications:
Experience leading combined software engineering and day-to-day operational work with accountability for application development, reliability, and operational excellence.
Experience managing globally distributed engineering teams, including onshore and offshore resources, across multiple application domains.
Experience owning 24x7 on-call, incident response, and escalation processes for business-critical systems, including driving post-incident reviews and corrective actions.
Experience establishing problem management practices focused on root cause analysis, recurring issue elimination, and continuous service improvement.
Experience implementing automation and AI-driven solutions to reduce manual operational effort, improve efficiency, and accelerate incident resolution.
Experience supporting supply chain, transportation, fulfillment, logistics, or customer delivery platforms in a large-scale enterprise environment.
Strong executive communication skills with the ability to translate complex technical issues into clear business impact, risk, and recovery plans.
Experience defining and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and service performance metrics across complex application ecosystems.
Proven ability to recruit, develop, mentor, and retain high-performing engineering talent while fostering a culture of accountability and continuous learning.
Experience developing technology roadmaps, driving quarterly planning, and aligning engineering priorities with business objectives.
Experience managing engineering capacity, operational workload, cloud consumption, and technology budgets in a cost-conscious environment.
Experience overseeing modern CI/CD pipelines, release management processes, and change governance practices to support reliable software delivery.
Experience driving system resiliency, disaster recovery readiness, capacity planning, and performance optimization for high-volume platforms.
Strong knowledge of observability platforms, including logging, metrics, tracing, and monitoring tools such as Datadog, Splunk, Grafana, Prometheus, New Relic, or Elastic.
Experience operating cloud-native applications on Google Cloud Platform, or Microsoft Azure, including container platforms such as Kubernetes.
Experience with Infrastructure as Code and configuration management tools such as Terraform, Ansible, Chef, or Puppet.
Experience supporting hybrid technology environments that include cloud services, on-premises platforms, vendor-supported applications, and relational databases.
Proficiency in at least one modern programming or scripting language such as Python, Java, Go, or Bash.
Experience partnering with software engineering, quality engineering, security, infrastructure, and business teams to improve application reliability and operational effectiveness.
Experience collaborating with security and compliance teams to meet regulatory and corporate policy requirements, including PCI-DSS and SOC 2.
Working knowledge of identity, access management, secrets management, credential lifecycle management, and security best practices in enterprise environments.
Minimum Education:
Preferred Education:
Minimum Years of Work Experience:
Preferred Years of Work Experience:
Minimum Leadership Experience:
Preferred Leadership Experience:
Certifications:
Competencies:
For California, Colorado, Connecticut, Rhode Island, Nevada, New York City, Ithaca (NY), Westchester County (NY), and Washington residents:
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”