For Employers

Lifted, an Upwork Company™

#127606 - Senior Software Engineer / SRE (Observability Focus)

Posted a month ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

The role involves designing, building, and maintaining software, APIs, and automation to support platform reliability and observability. You will also be responsible for monitoring containerized environments and driving proactive maintenance and performance improvements.

Job Description

We are seeking a Senior Software Engineer / SRE with a strong observability focus to support platform reliability, monitoring, and modernization initiatives. This role combines approximately 60–70% software engineering with 30–40% site reliability engineering and requires hands-on experience working with Kubernetes, cloud infrastructure, observability platforms, APIs, and operational automation.

 

The successful candidate will help build and operate observability capabilities across containerized and microservices-based environments while proactively improving platform reliability, scalability, and performance.

 

Enterprise experience strongly preferred.

 

Key Responsibilities:

 

- Design, build, and maintain software, APIs, integrations, and automation that support platform reliability and observability.

- Support monitoring, reliability, maintenance, and continuous improvement across internal platforms and systems.

- Work in Kubernetes-based environments across deployment, operations, and monitoring activities.

- Build and maintain observability solutions with a focus on Datadog.

- Configure dashboards, alerts, application performance monitoring, tracing, metrics, and logging.

- Monitor containerized and microservices-based applications.

- Integrate observability platforms and monitoring capabilities into AWS environments.

- Integrate observability capabilities into CI/CD pipelines and deployment processes.

- Automate monitoring and operational tasks through scripting, with Python preferred.

- Install and configure Datadog agents and integrations.

- Manage observability API keys and secure configurations.

- Manage user roles, permissions, and access controls within observability platforms.

- Lead proactive maintenance efforts and platform improvements.

- Drive improvements in reliability, scalability, performance, and operational efficiency.

Qualifications

Must-Have Skills:

 

- Strong proficiency in at least one of Python, JavaScript using Node.js, or Java.

- Hands-on experience designing, consuming, and implementing API integrations.

- Strong Kubernetes experience covering deployment, operations, and monitoring.

- Hands-on experience with Datadog or a comparable observability platform such as Prometheus or Grafana.

- Experience configuring dashboards, alerts, application performance monitoring, tracing, metrics, and logging.

- Experience monitoring containerized and microservices-based architectures.

- Hands-on AWS experience.

- Experience integrating observability tooling into cloud environments.

- Experience integrating observability capabilities into CI/CD pipelines.

- Ability to automate monitoring and operational work through scripting.

 

Nice-to-Have Skills:

 

- Strongly preferred experience owning and operating an internal engineering platform.

- Strongly preferred experience owning reliability, scalability, and performance outcomes.

- Strongly preferred experience proactively leading maintenance efforts and platform improvements rather than providing only reactive support.

- Familiarity with Go or Golang.

- Experience with New Relic, Dynatrace, Elastic, or Splunk Observability.

- Experience working across multiple observability and monitoring platforms.

Additional Information

Required Tools & Platforms:

 

- Python, JavaScript using Node.js, or Java

- Kubernetes

- Datadog, Prometheus, Grafana, or a comparable observability platform

- AWS

- CI/CD pipelines

- APIs and integration tooling

- Application performance monitoring, tracing, metrics, logging, dashboards, and alerting tools

 

Location, Time & Engagement:

 

- Remote contract opportunity.

- Candidates must be located in APAC.

- Ability to provide overlap with Japan Standard Time is preferred.

- Full-time allocation of 40 hours per week.

- Expected contract end date is March 31, 2027.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Docentes para el Grado en Animación 2D y 3D

Part Time Spain Software Development

Senior Software Engineer - Integrations & Partnerships

Full Time Ireland, United Kingdom Software Development

Senior Business Analyst – Workers Compensation

Full Time United States Software Development

Senior Java Engineer with GCP

Full Time Ukraine Software Development

Software Engineer Intern

Internship, Part Time Worldwide Software Development

Web Designer - WordPress & Elementor (WFH) | ZR_1323_JOB

Full Time Australia Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Software Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified