For Employers

ZENITH INFOTEK LLC

Site Reliability Engineering SRE Leader with AI Experience

Posted 21 days ago
10+ years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

The SRE Leader will design and implement scalable, high-performance enterprise applications while leveraging AI and Generative AI technologies to improve system reliability. They will collaborate with cross-functional teams to build automation solutions, manage CI/CD pipelines, and ensure operational excellence through rigorous monitoring and capacity planning.

Job Title: Site Reliability Engineering (SRE) Leader with AI Experience
Location: (Remote)
Duration: 9+ Months (Contract)
Interview: Video Interview


Job Summary
We are seeking an experienced Site Reliability Engineering (SRE) Leader with strong expertise in Java application development, cloud technologies, DevOps, and AI-driven solutions. The ideal candidate will lead initiatives focused on improving application reliability, scalability, automation, and operational excellence while leveraging modern AI and Generative AI technologies. This role requires close collaboration with engineering, infrastructure, and operations teams to build highly available, resilient, and high-performing enterprise applications.


Key Responsibilities
  • Design, develop, and implement scalable, reliable, and high-performance enterprise applications.
  • Collaborate with development, infrastructure, and operations teams to improve system reliability and availability.
  • Build and enhance automation solutions to streamline deployment, monitoring, and operational processes.
  • Develop and maintain CI/CD pipelines using Jenkins, CloudBees, or similar tools.
  • Monitor application health using observability platforms such as New Relic, Datadog, Dynatrace, and Splunk.
  • Analyze system performance, identify bottlenecks, and implement performance optimization strategies.
  • Design solutions for high availability, redundancy, auto-recovery, and fault tolerance.
  • Perform capacity planning and manage application performance against SLA/SLO objectives.
  • Lead Proof of Concepts (POCs) and successfully scale them into enterprise-wide implementations.
  • Implement best practices for application lifecycle management, monitoring, and continuous improvement.
  • Work with AWS cloud technologies to deploy and manage cloud-native applications.
  • Evaluate and integrate AI and Generative AI technologies into operational workflows.
  • Mentor engineering teams on SRE principles, DevOps practices, and reliability engineering.
  • Stay current with emerging technologies and recommend innovative solutions to improve platform stability and efficiency.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
  • 10+ years of experience in enterprise Java application development and runtime support.
  • Strong experience with:
    • Java
    • Spring Boot
    • Microservices Architecture
    • Java Enterprise Frameworks
  • Deep understanding of:
    • Site Reliability Engineering (SRE)
    • Reliability Engineering concepts
    • Automation
    • High Availability
    • Auto Recovery
    • Scalability
    • Redundancy
  • Experience with AWS Cloud services (AWS Certification preferred).
  • Hands-on experience with CI/CD tools such as Jenkins or CloudBees.
  • Strong knowledge of observability and monitoring tools including:
    • New Relic
    • Datadog
    • Dynatrace
    • Splunk
  • Experience with PCF (Pivotal Cloud Foundry), App Pilot, and Presto.
  • Strong understanding of application infrastructure, runtime environments, capacity planning, and SLA/SLO management.
  • Experience designing and scaling Proof of Concepts (POCs) into enterprise-grade solutions.
  • Familiarity with Agile methodologies and DevOps practices.
  • Excellent troubleshooting, analytical, and problem-solving skills.
Preferred Qualifications
  • Experience with Artificial Intelligence (AI) technologies.
  • Hands-on knowledge of Generative AI platforms and tools such as:
    • Google Cloud AI (Vertex AI/GHCP)
    • Claude
    • AWS Bedrock
  • Experience implementing AI-enabled operational automation.
  • AWS Professional or Specialty Certifications are a plus.
Technical Skills
  • Java
  • Spring Boot
  • Microservices
  • AWS Cloud
  • Jenkins
  • CloudBees
  • Splunk
  • Datadog
  • Dynatrace
  • New Relic
  • PCF (Pivotal Cloud Foundry)
  • App Pilot
  • Presto
  • CI/CD
  • DevOps
  • Site Reliability Engineering (SRE)
  • Capacity Planning
  • SLA/SLO
  • AI
  • Generative AI
  • AWS Bedrock
  • Claude
  • Google Cloud AI
Agile

This is a remote position.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Software Engineer II

Other Mexico 583K - 875K per year Software Development

Senior Business Analyst - Healthcare

Full Time United States $69400 - $99200 per year Software Development

Software Engineer II

Other Mexico 583K - 875K per year Software Development

Staff Clinical Informatics Data Architect

Full Time United States Software Development

Principal AI Solutions Architect

Full Time United States Software Development

Lead Engineer, Growth

Full Time United States $200K - $220K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Software Development

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified