Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will build and operate high-volume Kubernetes infrastructure, ensuring reliability for stateful, long-lived messaging connections. Responsibilities include managing observability, networking, disaster recovery, and participating in a 24/7 on-call rotation.
Use AI-assisted engineering where it improves delivery - including infrastructure tooling, automation, documentation and runbooks.
Build and operate Kubernetes infrastructure across two sites using infrastructure-as-code and automated delivery.
Design deployment mechanisms that allow long-lived connections to drain gracefully instead of being dropped during releases.
Take ownership of platform networking, including stable ingress and egress IPs, Layer 4 load balancing, TLS and connectivity coordinated with an external infrastructure provider.
Build and operate the observability stack: metrics, logs, dashboards and alerting that provide useful signal rather than noise.
Design monitoring around real production behaviour, including visibility into individual connections and failure modes.
Own certificate lifecycle, secrets management and platform access control.
Operate PostgreSQL in a highly available environment, including replication, failover, backups and verified restores.
Design and exercise backup and disaster-recovery procedures across two sites.
Work closely with engineers and QA on performance and production-scale load testing, investigating what fails first and why.
Support the platform during migration and production cutovers.
Take part in production incident response and from the steady-state phase, a 24×7 on-call rotation.
Experience with telecom, carrier, messaging or other environments built around persistent network connections.
Knowledge of SMPP, SIP, SS7 or similar telecom protocols.
Strong production experience operating Kubernetes on self-managed, on-premise or similarly infrastructure-heavy environments - not only managed cloud services.
Experience with stateful, long-lived TCP workloads on Kubernetes, including connection draining, stable ingress/egress, Layer 4 load balancing and deployment behaviour.
Strong Linux and networking fundamentals. You are comfortable diagnosing problems involving routing, NAT, firewalls, MTU, TLS or packet-level behaviour.
Practical troubleshooting experience with tools such as tcpdump and production network diagnostics.
Experience designing monitoring and alerting, not only maintaining dashboards somebody else created. You understand what deserves to wake a human up, and what does not.
Strong infrastructure-as-code and CI/CD experience for containerised systems.
Practical PostgreSQL operations experience, including replication, failover, backup and importantly - verified restore.
Ability to work directly with engineers from external infrastructure and network providers.
Professional English.
Real production on-call and incident-response experience.
Experience configuring or troubleshooting IPsec connectivity with external parties.
Experience designing or operating multi-site active-active or active-passive environments.
Hands-on disaster-recovery exercises rather than DR plans that existed only on paper.
Performance engineering experience, including Linux kernel or network tuning for high connection counts.
Security hardening experience, including CIS-style benchmarks, vulnerability management or software supply-chain practices.
Experience being the first SRE or Platform Engineer on a system and defining how it should be operated.
Build it and run it: you won’t inherit an infrastructure somebody else designed or throw your work over the wall after launch. You’ll help build the platform and remain close to it in production.
A genuinely difficult reliability problem: long-lived protocol traffic, stateful connections, two-site infrastructure and production traffic at telecom scale.
Influence architecture from day one: deployment, observability and operability are design constraints here, not tasks postponed until after development.
Greenfield infrastructure: you’ll have real influence over how the Kubernetes platform, delivery pipelines, monitoring and operational practices are created.
Small senior team: short decision paths and direct collaboration with engineers and the solution architect.
Production ownership: you’ll follow the system through build, migration, hypercare and steady-state operation instead of disappearing after implementation.
Modern engineering environment: automation and AI-assisted tooling are used where they genuinely improve engineering and operational work.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 221,977+ Jobs in DevOps Engineer
Answer easy questions
221,977+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”