For Employers

AgileEngine

Senior Cloud/DevOps Engineer ID92208

Posted a day ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

The engineer will operate and improve AWS infrastructure and Kubernetes-hosted workloads for an enterprise data platform. They will also lead incident response, perform root-cause analysis, and mentor team members while ensuring system reliability.

AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

ABOUT THE ROLE
We are looking for a Senior Cloud/DevOps Engineer to operate and improve the cloud infrastructure and reliability layers behind an enterprise data platform in a regulated healthcare environment.

The mandatory requirements are 5+ years of experience in Cloud Engineering, DevOps, or Site Reliability Engineering, advanced hands-on experience with AWS EKS and Kubernetes, experience administering and troubleshooting Argo Workflows, and strong English communication skills.


MUST HAVES
- 5+ years of professional experience in Cloud Engineering, DevOps or Site Reliability Engineering.
- Strong hands-on experience operating AWS infrastructure in production environments.
- Advanced experience with Kubernetes and Amazon EKS, including workload operations, troubleshooting, access, observability, capacity, and reliability.
- Hands-on experience administering and troubleshooting Argo Workflows or comparable workflow orchestration platforms.
- Strong Infrastructure as Code experience with Terraform and source-controlled infrastructure practices.
- Experience building, hardening, and supporting CI/CD pipelines and production release processes.
- Strong experience with monitoring, logging, alerting, and incident-routing tools such as Splunk, PagerDuty, Opsgenie, or comparable platforms.
- Demonstrated ability to lead complex incident resolution, perform root-cause analysis, and translate findings into preventive improvements.
- Proficiency in automation and scripting using Python, Shell, Bash, or similar languages.
- Ability to make well-reasoned technical decisions, identify tradeoffs, estimate work, and drive improvements across a complex platform.
- Experience mentoring engineers and collaborating effectively with Data Engineering, Security, Governance, Analytics, and business stakeholders.
- Strong written and verbal English communication skills, with the ability to work directly with client stakeholders.
- Availability to work within the LatAm service window of approximately 9:00 AM to 6:00 PM Eastern Time and participate in an agreed on-call rotation.

NICE TO HAVES
- Experience supporting data-platform infrastructure involving Snowflake, dbt, Fivetran, HVR, Tableau Cloud, or custom ingestion pipelines.
- Familiarity with data-specific observability platforms such as SYNQ.
- Experience modernizing or migrating legacy orchestration and ingestion solutions such as Boomi or AWS Data Pipeline.
- Experience with service-management and change-control tools such as Freshservice and Jira.
- Experience operating in healthcare, life sciences, financial services, or another regulated environment.
- Familiarity with HIPAA, GDPR, FDA-related controls, least-privilege access, separation of duties, and audit-ready operational practices.

WHAT YOU WILL DO
- Provide senior technical ownership for the Cloud / DevOps service tower during the LatAm coverage window, including day-to-day operations, complex troubleshooting, and L2/L3 escalation.
- Operate, maintain, and improve AWS infrastructure supporting the Data Platform, including Amazon EKS, S3, EventBridge, SQS, API Gateway, Lambda, and related services.
- Administer Kubernetes-hosted workloads and Argo Workflows, including deployment, scheduling, monitoring, troubleshooting, capacity management, resiliency, and recovery.
- Define and improve standards for Infrastructure as Code, configuration management, CI/CD, release execution, rollback, and environment consistency, primarily using Terraform and Git-based delivery practices.
- Lead the consolidation and improvement of observability across infrastructure and data workloads, linking alerts to operational evidence from Argo, dbt, Snowflake, and supporting runbooks.
- Improve alert routing and escalation workflows across tools such as Splunk, Opsgenie, PagerDuty, Microsoft Teams, and data-specific observability platforms.
- Lead or support major incident response, root-cause analysis, post-incident reviews, and corrective actions, with clear communication to technical and service stakeholders.
- Design and implement reliability improvements such as selective auto-remediation, dependency-aware alert correlation, impact analysis, and automation of repetitive operational work.
- Track and contribute to service metrics including availability, SLA compliance, alert volumes, workflow reliability, deployment outcomes, and mean time to restore service.
- Apply disciplined change-management, access-control, secrets-management, auditability, and documentation practices appropriate for a HIPAA-, GDPR-, and FDA-regulated environment.
- Create and maintain runbooks, operating procedures, architecture context, recovery procedures, and knowledge-transfer materials.
- Mentor Middle-level engineers, review technical work, improve team practices, and promote consistent execution across the distributed team.
- Participate in the Cloud / DevOps on-call rotation for critical incidents outside staffed service hours.

PERKS AND BENEFITS
- Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
- Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
- Exciting projects: Modern solutions with Fortune 500 and top product companies.
- Flextime: Flexible schedule with remote and office options.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Implementation Engineer

Full Time Software Development

Desenvolvedor Full Stack Sr (Python e React)

Contract Software Development

Director of Quality Assurance - Student Experience and New School Integration (Immediate Opening)

Full Time $108K - $128K per year Software Development

Bilingual Customer Success Operations Analyst Fleet Analytics & Digital Solutions

Freelance Software Development

Bilingual Software Tester (CYARA Contact Centre and AWS Connect)

Freelance 45 - 65 per hour Software Development

Software Developer (Co-op Opportunities)

Internship 30 per hour Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 216,181+ Jobs in Cloud DevOps Engineer

Answer easy questions

Answer easy questions

216,181+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified