You will design, build, and operate cloud-native Kubernetes platforms while collaborating directly with client engineering teams. You are responsible for maintaining platform reliability, implementing infrastructure as code, and establishing self-service paths for delivery engineers.
EggAI · Full-time · Remote (EU only)
The Role
A Senior Platform Engineer with deep, hands-on experience designing and operating cloud-native platforms in production.
You will work in a team of 2 to 6 engineers on a client project — one project at a time, so you go deep rather than spreading yourself across accounts. Clients come to us because we work at the forefront of AI in the enterprise; you build and run the platform their AI systems actually land on, alongside their own infrastructure engineers.
What You'll Do
Design, build and operate the Kubernetes platform for the engagement, defined in Terraform — you own it in production, not just at handover
Lead architecture discussions for your part of the platform, with the tech lead supporting you
Take ownership: pick up the ambiguous problem, chase down the answer, and say so when something is going wrong
Build the self-service paths the delivery engineers use, so they aren't blocked on you
Own reliability for what you run — SLOs, observability, incident response, cost
Work directly with the client's engineers, in a normal team rhythm: daily standups, planning, retrospectives, code and design review
What We're Looking For
5+ years in infrastructure / platform / SRE roles, at least 3 of them running production Kubernetes — cluster design and operation, workload lifecycle, RBAC, secrets, multi-tenancy
Infrastructure as Code in production: Terraform (primary), plus Helm or Kustomize — owned over time, not written once
A major cloud in depth (AWS, GCP or Azure): networking, compute, managed databases, IAM and least-privilege, and a habit of watching what it costs
CI/CD and GitOps: building pipelines and running ArgoCD, Flux or equivalent
Observability and reliability: metrics, logs and tracing (Prometheus/Grafana, OpenTelemetry or similar), SLOs, and on-call incident response
Scripting and tooling: Bash, plus Python or Go
Security as a first-class concern: secrets management, supply-chain and vulnerability scanning
Client-facing ability: can hold a technical conversation with a client's architects and engineers
Desired qualities
Platform-thinking mindset: builds for internal customers, not just operations tasks
Pragmatic about trade-offs between reliability, cost, and engineering velocity
Thinks about blast radius before shipping
Clear written communication — design docs, trade-offs, post-incident reviews
Nice to have
Service mesh, multi-cluster or multi-region experience
Consulting or professional-services background
AI/ML or agentic workloads in production
EU sovereign cloud, data residency or public sector infrastructure experience
About EggAI
EggAI is an enterprise-focused generative AI company on a mission to help large organizations move AI solutions from prototyping into production. We specialize in building safe, reliable, and scalable AI systems that deliver real business impact.
We work at the cutting edge of AI technology, implementing agentic systems, RAG (Retrieval-Augmented Generation) architectures, and autonomous AI agents that scale from task automation to workforce automation. Our proprietary frameworks—including the EggAI Meta Framework for agentic systems and EggAI Quality Flow for governance—power AI transformations at enterprise scale.
Based in Munich, Germany, we work with enterprise clients to build AI capability, deliver production-ready systems, and establish quality-controlled AI operations.
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”