Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
Type: Contract, per-project.
Location: Remote - within LATAM time zones (GMT-3 to GMT-5) - Remote with meetings across U.S. time zones
Availability: Contractor (40 hours per week)
We are looking for a Platform Observability Engineer to help onboard application services onto our standard Grafana observability stack and establish observability configurations as code.
This role focuses on observability, infrastructure automation, monitoring and alerting, and working closely with application teams to define and implement useful operational signals.
Onboard application services onto the observability stack, including Grafana, Loki for logs, Mimir for metrics, Tempo for traces, and OpenTelemetry collectors for instrumentation.
Work with application teams to instrument their services, identify meaningful observability signals, and implement dashboards and alerts that support their operational needs.
Implement dashboards, alert rules, and collector configurations as code using Terraform and pull requests.
Migrate alerting from legacy monitoring tools to the standard observability stack, validating parity before retiring existing checks.
Support and maintain observability configurations across application services.
Document each service and its observability setup as it is onboarded.
2–5 years of experience in Platform Engineering, Site Reliability Engineering (SRE), Cloud Engineering, DevOps, or a similar role.
Hands-on experience with Terraform, including writing modules, managing state, and working through a pull-request-based workflow.
Production experience with observability tooling, including Grafana and Grafana Alerting, plus at least one of the following:
Prometheus.
Loki.
Tempo.
OpenTelemetry Collectors.
Working knowledge of AWS, particularly ECS.
Experience with Python for tooling and automation.
Linux fundamentals.
Clear written English.
Ability to work effectively with application teams that are not observability specialists and help them adopt standardized observability practices.
Experience with OpenTelemetry instrumentation inside a real application codebase, beyond collector configuration.
Experience with Mimir or another Prometheus-compatible long-term metrics storage solution.
Experience with Nagios or comparable legacy monitoring tools, including the ability to understand and migrate existing check definitions.
Experience migrating alerting from legacy monitoring systems to modern observability platforms.
Experience with GCP or OCI.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Development
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!