Match your resume skills with our AI powered skill match!
Questions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will own the platform architecture, security, and scalability for AI agents across Azure and AWS environments. You are responsible for managing Kubernetes clusters, CI/CD pipelines, and ensuring the platform remains secure, observable, and cost-efficient.
As Senior Platform Engineer, you own the platform that Workist's AI agents run on: two clouds, the Kubernetes clusters for production, staging and ML workloads, and the path from merge request to production for every service we ship. You decide how this platform evolves, you keep it secure, observable and cost-efficient, and you extend it as the product grows. Most recently that meant an LLM gateway with region failover in front of our Azure OpenAI deployments, and performing GPU capacity planning for our own models.
You set your own roadmap. Infrastructure at Workist is run as quarterly themes that you propose, make the case for and deliver, and roughly a third of your time goes to whatever the week brings: an incident, a pentest finding, a developer whose deployment is stuck and needs a second pair of eyes. You report to our CTO and are the voice of the platform in engineering decisions.
Everything runs on Kubernetes (Azure AKS), deployed with Helm and GitOps/Flux. All infrastructure is Terraform, applied through CI.
CI/CD for every service on GitLab
Managed services wherever they keep life simple: Postgres, OpenSearch, Redis, blob storage. We self-host only where it clearly pays off.
Two clouds: Azure as our primary cloud, AWS for search and mail ingestion.
Python everywhere: read and fix application code when that is where the fix belongs.
Security is routine: automated scanning in every pipeline, regular external pentests, quarterly backup and disaster-recovery tests.
Own the architecture of our cloud environments across Azure and AWS: how subscriptions, networks, identities and environments are structured and stay isolated
Design and deliver the platform capabilities the product needs next, from GPU capacity and an LLM gateway to new environments
Set the standard for how we run Kubernetes and our managed data services: capacity planned ahead of growth, changes and upgrades rolled out safely and routinely, costs kept in check
Own the security posture end to end: least-privilege access, network isolation, findings from scanners and pentests, and a patch cadence that runs itself. Be a technical counterpart for audits and security questionnaires
Own observability and alerting so that every alert is worth acting on, and lead the response to platform incidents including the follow-up fixes
Run quarterly backup/restore and disaster-recovery tests, and close what they reveal
Own the path from merge request to production: fast builds, reliable deployments, environments and access on demand
Build the infrastructure behind our AI models: the LLM gateway, model deployments with region failover, GPU capacity, and whatever the next model needs
We're looking for a professional team player who thrives in an environment where we share knowledge openly and push each other to deliver the best possible outcome. What you bring:
You have run production infrastructure for years, ideally as the person who owned it end to end rather than one of many in a large ops team
Kubernetes and Terraform in production: cluster upgrades, Helm, GitOps, and infrastructure as code for a whole cloud estate
Production experience on Azure or AWS. Most of our estate is Azure, so an AWS background works if you are willing to go deep on Azure.
Active Python skills: you read and debug a Python codebase and fix the application when the application is what is broken
Security as a habit: you have run patch management, handled pentest findings and designed least-privilege access, and you prioritise findings pragmatically
You write clearly: runbooks, incident write-ups and architecture decisions that others can follow
Fluent English (engineering runs in English, so German is not required, though it can help in some customer and vendor conversations)
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 216,786+ Jobs in Platform Engineer
Answer easy questions
216,786+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”