Please mention DailyRemote when applying
See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewUpload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will evolve the observability stack for logs, metrics, and traces while ensuring reliable telemetry across cloud and on-premise dataplanes. Additionally, you will lead incident response, define SLOs, and optimize telemetry costs to maintain high system availability.
At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. Seniority shows up in your scope, level, and compensation, not in titles.
In this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments — so observability and reliability have to work the same way everywhere.
Evolve our observability stack for logs, metrics, traces, and alerting.
Make sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane.
Bring telemetry into customer clusters within a model where agents only make outbound connections.
Detect drift between the desired state and what's actually running in each environment.
Monitor the health of our deployment and runtime agents.
Provide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference.
Define SLOs, lead incident response and postmortems, and reduce MTTR — including when a fix requires coordinating with the customer.
Reduce telemetry cost: less redundant data, more useful signal.
99.9% serving availability, with incidents trending down.
MTTR, including on-premise incidents.
Near-zero drift between desired and actual state.
All agents active and reporting, across every dataplane.
Deep experience with OpenTelemetry and observability backends.
Hands-on practice with SLOs, error budgets, actionable alerting, and incident management.
Strong experience with Kubernetes and infrastructure as code (Terraform / Helm ).
Experience operating software in environments you don't fully control.
Production-quality code and reviews, and a willingness to operate what you build.
Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity).
GCP or GKE, AWS or EKS.
ML multi-node/multi-cluster workloads in production.
Financial services or regulated environments.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 218,630+ Jobs in Others
Answer easy questions
218,630+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”