Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
Questions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will design, build, and operate cloud infrastructure while automating manual processes to ensure system reliability and scalability. Additionally, you will manage deployment pipelines, maintain observability tools, and mentor engineers to improve delivery velocity.
MetaRouter is a customer-data streaming platform for large, security-conscious enterprises. The MetaRouter platform dramatically reduces latency and bloat by providing server-side integration with third-party marketing, analytics, and data storage/transport tools. By enabling customers to access our SaaS tool, access a private PaaS architecture, or deploy fully within their own private cloud, customers can dramatically simplify and centralize their customer-data pipelines and maintain full control over security and compliance.
The MetaRouter platform, serving as a full-spectrum customer data collection, modification, and delivery platform, is a collection of many, varied microservices. Our client and server-side ingestion and identity libraries, ETL applications, and configuration and monitoring UIs serve to give data teams control over the shape and substance of their customer-behavior data, all while powering the flexibility and freedom to be creative with their architecture.
As a Senior Site Reliability Engineer, you own significant pieces of our infrastructure and operational tooling end to end. You will automate what is manual, instrument what is opaque, and make our deployments repeatable across a growing number of isolated customer environments.
We run dedicated, private environments per customer, so the interesting problems here are about repeatability, automation, and observability at scale. Experience with these patterns matters more than familiarity with any particular cloud, orchestrator, or monitoring vendor.
Design, build, and operate the cloud infrastructure that supports our applications and internal operations, from provisioning through decommissioning.
Own deployment automation and release tooling, and improve the safety and speed of getting changes to production.
Manage upgrades across infrastructure and the supporting software our applications depend on.
Build and maintain dashboards, logs, metrics, and alerting so problems are caught early and alerts mean something.
Troubleshoot infrastructure and application issues in production, and drive fixes through to resolution.
Scale systems sustainably through automation, and push for changes that improve both reliability and delivery velocity.
Handle identity and access management and single sign-on across platforms and services.
Ensure infrastructure and applications meet compliance requirements.
Work with customers to determine and implement custom infrastructure requirements.
Participate in code reviews to ensure infrastructure, applications, and supporting services follow best practice.
Improve and maintain infrastructure and process documentation.
Mentor engineers earlier in their careers and pair with product teams to spread reliability practice.
Participate in the on-call rotation.
6+ years in an SRE, DevOps, or platform engineering role operating production systems.
Hands-on experience with at least one major public cloud provider.
Solid experience configuring, maintaining, and troubleshooting container orchestration clusters.
Working fluency with infrastructure as code, configuration management, containerization, version control, and CI/CD tooling.
Strong scripting and automation skills, plus the ability to debug and optimize application code when needed.
Experience building dashboards, metrics, and alerts in a modern observability platform, including its query language.
Expertise troubleshooting large-scale distributed systems, with a systematic problem-solving approach.
Solid understanding of Unix/Linux operating systems and networking fundamentals.
Effective written and verbal communication, and comfort prioritizing a wide variety of tasks in a fast-paced environment.
Familiarity with agile methodologies.
Job Type: Full Time Location: Fully Remote (US)
Health / Dental / Vision insurance
401(k)
Unlimited vacation policy
Fully remote (US)
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 204,268+ Jobs in Site Reliability Engineer
Answer easy questions
204,268+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”