Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
The Senior Infrastructure Operations Engineer owns the global design, administration, and continuous improvement of compute, storage, and cloud infrastructure. This role serves as the primary technical escalation point and leads infrastructure projects while mentoring team members.
The Senior Infrastructure Operations Engineer owns the design, administration, and continuous improvement of KLDiscovery's compute, storage, and cloud infrastructure globally. This role takes end-to-end technical ownership of the physical server, virtualization, block/file/object storage, enterprise backup, and Azure IaaS environments underpinning KLDiscovery's production client systems, and serves as the primary technical escalation point and design authority within the Compute & Storage team. The Senior Infrastructure Operations Engineer leads infrastructure design decisions, drives standards development, partners with Enterprise Architecture and IT Security on architecture and compliance, and provides mentoring and technical direction to Infrastructure Operations Engineers. This role participates in 24x7x365 on-call rotation.
Key Responsibilities
Compute, Virtualization & Storage Architecture:
Hold end-to-end technical ownership of KLDiscovery’s physical server and virtualized compute environments; define and enforce configuration standards, capacity thresholds, and operational procedures; lead the design of significant compute changes, cluster expansions, and platform upgrades; review and approve significant configuration changes before implementation
Own the design standards, capacity planning, performance management, and operational procedures for block, file, and object storage environments globally; monitor performance and capacity; identify constraints and drive remediation before they impact production; engage vendor support directly for complex issues
Backup, Recovery & Azure IaaS:
Own the enterprise backup platform and recovery strategy — job standards, recovery validation procedures, and RPO/RTO alignment; ensure recovery procedures are documented, tested on a defined schedule, and executable by any team member; manage the formal backup coverage request process for consuming teams
Own the administration and governance of Azure IaaS infrastructure; define Azure infrastructure standards in conjunction with Enterprise Architecture; monitor and govern cloud costs; surface anomalies and optimization opportunities proactively; partner with Enterprise Architecture on hybrid cloud architecture direction and roadmap input
OS Standards, Patching & Provisioning:
Own Windows Server and Linux configuration standards, patching cadence, and hardening baselines across all infrastructure globally; own the server patching function for all server infrastructure — end-user endpoints are excluded and owned by Enterprise Platforms
Define and own the server build runbook, provisioning standards, sizing guidelines, and handoff checklists in conjunction with Enterprise Architecture; ensure the provisioning process delivers a configured, network-connected OS at the domain-join boundary with sufficient fidelity to eliminate rework at handoff
Security, Observability & Escalation:
Embed security controls into infrastructure design from inception — network segmentation, least-privilege access, encryption at rest and in transit, and audit logging — in conjunction with Enterprise Architecture and IT Security; support audit and compliance reviews as needed; maintain working familiarity with applicable security and compliance frameworks (ISO 27001, CIS Controls) as applied to infrastructure configuration and access management
Partner with the Automation & Observability team to ensure all owned infrastructure is covered by monitoring and alerting; serve as the primary escalation point within the Compute & Storage team for complex or time-critical infrastructure incidents; participate in the 24x7x365 on-call rotation; conduct root cause analysis for significant incidents, lead post-incident reviews and blameless retrospectives, and drive systemic remediation
Continual Service Improvement & Automation:
Evaluate existing infrastructure for improvement, consolidation, and modernization opportunities; make specific, cost-weighted technical recommendations to management for platform upgrades, replacements, or cloud migrations
Partner with the Automation & Observability team to identify, prioritize, and define requirements for infrastructure automation; serve as the infrastructure subject matter expert in the development of IaC solutions; own the execution, validation, and operational adoption of approved automated solutions within the Compute & Storage environment.
Own and govern infrastructure performance KPIs — availability, capacity utilization, incident response SLAs, and patching compliance; proactively surface trends and recommend corrective actions
Strategy, Cross-Team Coordination, Projects & Documentation:
Serve as the primary technical liaison to Database, Enterprise Platforms, Automation & Observability, Networking, and DC Operations; resolve cross-team disagreements with peer team leads; escalate unresolved issues per the established 48-hour escalation model
Translate architectural designs from Enterprise Architecture into operational infrastructure standards and procedures; provide input into overall infrastructure roadmap planning and technology lifecycle decisions
Lead technical delivery of infrastructure projects from inception to completion with measurable milestones and managed scope; engage and manage third-party vendors on complex infrastructure issues, support renewals, and service escalations
Own the documentation standard for the Compute & Storage team; drive creation and maintenance of the server build runbook, capacity and resource catalog, infrastructure configuration standards, provisioning handoff checklist, backup coverage map, and recovery runbooks
Decision Scope & Accountability: Authorized to make independent infrastructure configuration and architectural decisions within defined scope and standards. Exercises discretion on decisions with broader impact — consults the Manager, IT - Compute & Storage before committing to changes affecting cross-team dependencies, security posture, Azure spend, or architectural direction.
Budgetary Awareness: Weighs cost into all infrastructure recommendations — hardware, Azure consumption, licensing, and vendor renewals. Surfaces cost considerations and optimization opportunities proactively to the Manager, IT - Compute & Storage. Does not independently commit spend.
Mentoring & Development: Provides significant mentoring and coaching to Infrastructure Operations Engineers. Actively contributes to cross-training with Database, Automation & Observability, Enterprise Platforms, and DC Operations teams.
Skills & Qualifications
6+ years in infrastructure engineering with hands-on ownership across compute, storage, and cloud in a production enterprise environment; prior senior or lead individual contributor experience preferred
Expert-level administration of VMware vSphere and Nutanix in enterprise production environments; experience with cluster design, capacity planning, and lifecycle management
Deep experience with block and file storage technologies — RAID, SAN (Fibre Channel or iSCSI), NAS protocols, and enterprise storage array administration; Hitachi or equivalent required
Expert-level administration of Veeam or equivalent enterprise backup platform — architecture, job design, recovery validation, and RPO/RTO management
Strong working knowledge of Azure IaaS architecture — virtual machines, managed disks, virtual networking, and cloud cost governance; hybrid cloud experience required
Expert-level Windows Server administration — Active Directory, DNS, DHCP, Group Policy, clustering, and server hardening
Strong Linux (Ubuntu) administration — installation, configuration, patching, hardening, and troubleshooting in a production environment
Advanced PowerShell and/or Bash scripting
Working knowledge of IaC concepts and tooling (Ansible, Terraform, or equivalent) sufficient to define infrastructure requirements, review IaC solutions developed by the Automation & Observability team, and execute and validate automated workflows in production
Working familiarity with security and compliance frameworks (ISO 27001, CIS Controls) as applied to infrastructure configuration and access management; CompTIA Security+ or equivalent understanding expected
Familiarity with container technologies (Docker, Kubernetes) and their infrastructure dependencies in an enterprise environment
Proficient in ITSM processes and ITIL-based Incident, Problem, Change, and Capacity Management
Experience delivering and owning technical projects end-to-end, including vendor management and stakeholder communication
Effective communication with peers, management, vendors, and internal customers at all levels; ability to convey complex technical concepts to non-technical stakeholders
Experience supporting 24x7 global production environments; on-call availability required
Education: Bachelor’s degree in computer science, Information Technology, or equivalent experience
Preferred: ITIL V3/4; VMware VCP, Nutanix NCP, or Azure Administrator/Solutions Architect (AZ-104/AZ-305); Microsoft Certified: Azure Administrator Associate; Veeam certification; CompTIA Security+ or CISSP; prior experience in eDiscovery, legal technology, or a similarly regulated environment
Driving Career Growth, Benefit Excellence: The KLD Advantage
At KLD we invest in employees and their families by placing their wellbeing first. We offer competitive total compensation that includes base pay, bonus potential, inclusive benefits, wellness programs, and perks. We use market and industry data to inform pay decisions while considering geography and labor markets, individual experience, and business needs. India compensation is based upon the local competitive market.
Paid time off, that offers various time off options to help employees maintain a work-life balance, such as Casual, Earned, Sick, Special Leave, and Holidays
Ongoing learning and development, a focus on continuous professional development through various training and education reimbursement programs
A diverse and inclusive workplace where we all learn, grow, and achieve the greatest heights…together
A surrounding team of mission-driven individuals who genuinely love what they do
Free, fun, interactive and incentivized global wellness program that promotes the wellbeing of our employees
India compensation is based upon the local competitive market
Who We Are
KLDiscovery provides technology-enabled services and software to help law firms, corporations, and government agencies solve complex data challenges. With offices in 26 locations across 17 countries, KLDiscovery is a global leader in delivering best-in-class data management, information governance, and eDiscovery solutions to support the litigation, regulatory compliance, and internal investigation needs of clients. Our Nebula Ecosystem provides powerful end-to-end eDiscovery and enterprise-grade information governance. Through its global Ontrack data recovery business, KLDiscovery delivers world-class data recovery, disaster recovery, email extraction and restoration, data destruction, and tape management.
We Provide Equal Employment Opportunity.
At KLDiscovery we believe that inclusion and diversity make us stronger. We are committed to fostering an inclusive environment for all employees that enhances wellbeing and belonging. We welcome and celebrate individuals of all backgrounds, experiences, and perspectives.
We do not discriminate on the basis of race, color, religion, gender, pregnancy, gender identity, sexual orientation, national origin, age, disability, genetic information, veteran status, or any other protected status. We are happy to support you with any accommodation request at any stage in our hiring process
#LI-DNI
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Development
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”