See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewQuestions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
The Principal Engineer will define the technical direction and multi-year roadmap for Atlassian's Kubernetes, compute, and AI/ML infrastructure platforms. They will lead cross-team initiatives, design scalable control planes, and mentor senior engineers to ensure platform reliability and security.
Atlassian's mission "to unleash the potential of every team" is the guiding light behind what we do. Our products, including Jira, Confluence and Bitbucket, help teams everywhere work better together, from NASA to Cochlear.
Our office is in Bengaluru, but eligible candidates can work remotely across India: from home, from an office, or somewhere in between. We call this TEAM Anywhere.
The KubePlat group in Core Engineering owns Atlassian's entire Kubernetes ecosystem. Every Atlassian cloud product, including Jira, Confluence, Loom and Rovo, runs on the clusters, platform services and deployment systems we build and operate. Our mission is to provide Atlassian a secure, reliable, cost-efficient and globally distributed compute platform, as the foundation for building great products.
What we own:
Kubernetes infrastructure: hundreds of production clusters (EKS and GKE) across AWS, GCP and regulated cloud environments.
Fleet management: automated, cloud-agnostic lifecycle management for clusters and platform components, with safe progressive rollouts.
Kubernetes Platform-as-a-Service: an opinionated, multi-tenant platform for running Atlassian's microservices.
Cloud resource management: self-serve, Kubernetes-driven provisioning of the cloud resources that services depend on.
Secure sandboxed compute: strongly isolated environments for running untrusted code and AI-agent workloads.
Software delivery: artifact management, deployment orchestration and automated verification for thousands of deployments every day.
AI/ML infrastructure: the GPU compute, high-bandwidth networking and parallel filesystem ecosystem on Kubernetes that powers model training and AI inference behind Atlassian's AI features.
We are in the middle of a major transformation: moving to a multi-cloud, cell-based architecture, meeting FedRAMP and data-residency requirements, and building a reliable, cloud-native foundation for AI training and inference at scale.
As a Principal Engineer, you will set the technical direction for Atlassian's Kubernetes, compute and AI/ML infrastructure platform. You will lead engineers across several teams on our hardest problems, and take large, cross-team initiatives from design to launch.
Define the architecture and multi-year roadmap for a multi-cloud, cell-based Kubernetes fleet that scales to thousands of clusters.
Design Kubernetes-native control planes (controllers, operators, CRDs and policy) that give engineers self-serve compute, networking and cloud resources.
Raise platform reliability and security: reduce blast radius, build safe progressive rollouts, define SLOs, isolate workloads and support compliance (for example FedRAMP).
Lead compute efficiency through autoscaling, bin packing, capacity planning and cost optimisation.
Build the AI/ML infrastructure ecosystem on Kubernetes: GPU compute at scale, high-bandwidth, low-latency networking, and high-performance parallel filesystems for training and inference.
Influence engineering and product leaders globally, set architectural standards, and mentor senior engineers.
10+ years building and operating large-scale cloud infrastructure or platform-as-a-service systems.
Deep, hands-on Kubernetes expertise from running it in production at scale: internals, networking, multi-tenancy, and building operators and CRDs.
Strong experience with AWS and/or GCP (EKS/GKE) and distributed systems design.
Strong programming skills in one or more languages such as Go, Python, Java or similar.
Platform engineering experience with Infrastructure as Code, GitOps/CD (for example Terraform, Crossplane, ArgoCD) and observability.
A track record of leading technical direction across teams, with excellent communication and mentoring skills.
Experience with Kubernetes fleet management, cell-based architectures, sandboxing (gVisor, Firecracker) or regulated environments.
Experience providing AI/ML infrastructure: GPU clusters, high-bandwidth networking (for example RDMA, InfiniBand, AWS EFA, GPUDirect) or parallel filesystems (for example Lustre, Amazon FSx for Lustre, GCP Parallelstore).
Contributions to CNCF open-source projects.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 219,047+ Jobs in Principal Software Engineer
Answer easy questions
219,047+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”