The SRE Leader will design and implement scalable, high-performance enterprise applications while leveraging AI and Generative AI technologies to improve system reliability. They will collaborate with cross-functional teams to build automation solutions, manage CI/CD pipelines, and ensure operational excellence through rigorous monitoring and capacity planning.
Job Title: Site Reliability Engineering (SRE) Leader with AI Experience
Location: (Remote) Duration: 9+ Months (Contract) Interview: Video Interview
Job Summary
We are seeking an experienced Site Reliability Engineering (SRE) Leader with strong expertise in Java application development, cloud technologies, DevOps, and AI-driven solutions. The ideal candidate will lead initiatives focused on improving application reliability, scalability, automation, and operational excellence while leveraging modern AI and Generative AI technologies. This role requires close collaboration with engineering, infrastructure, and operations teams to build highly available, resilient, and high-performing enterprise applications.
Key Responsibilities
Design, develop, and implement scalable, reliable, and high-performance enterprise applications.
Collaborate with development, infrastructure, and operations teams to improve system reliability and availability.
Build and enhance automation solutions to streamline deployment, monitoring, and operational processes.
Develop and maintain CI/CD pipelines using Jenkins, CloudBees, or similar tools.
Monitor application health using observability platforms such as New Relic, Datadog, Dynatrace, and Splunk.
Analyze system performance, identify bottlenecks, and implement performance optimization strategies.
Design solutions for high availability, redundancy, auto-recovery, and fault tolerance.
Perform capacity planning and manage application performance against SLA/SLO objectives.
Lead Proof of Concepts (POCs) and successfully scale them into enterprise-wide implementations.
Implement best practices for application lifecycle management, monitoring, and continuous improvement.
Work with AWS cloud technologies to deploy and manage cloud-native applications.
Evaluate and integrate AI and Generative AI technologies into operational workflows.
Mentor engineering teams on SRE principles, DevOps practices, and reliability engineering.
Stay current with emerging technologies and recommend innovative solutions to improve platform stability and efficiency.
Required Qualifications
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
10+ years of experience in enterprise Java application development and runtime support.
Strong experience with:
Java
Spring Boot
Microservices Architecture
Java Enterprise Frameworks
Deep understanding of:
Site Reliability Engineering (SRE)
Reliability Engineering concepts
Automation
High Availability
Auto Recovery
Scalability
Redundancy
Experience with AWS Cloud services (AWS Certification preferred).
Hands-on experience with CI/CD tools such as Jenkins or CloudBees.
Strong knowledge of observability and monitoring tools including:
New Relic
Datadog
Dynatrace
Splunk
Experience with PCF (Pivotal Cloud Foundry), App Pilot, and Presto.
Strong understanding of application infrastructure, runtime environments, capacity planning, and SLA/SLO management.
Experience designing and scaling Proof of Concepts (POCs) into enterprise-grade solutions.
Familiarity with Agile methodologies and DevOps practices.
Excellent troubleshooting, analytical, and problem-solving skills.
Preferred Qualifications
Experience with Artificial Intelligence (AI) technologies.
Hands-on knowledge of Generative AI platforms and tools such as:
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”