Design and build a distributed control plane to run AI workloads across multiple clouds and on-prem infrastructure. Set the technical direction for the system to ensure it is robust, performant, and user-friendly.
About SkyPilot
SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."
SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.
The role
SkyPilot's core is a distributed control plane that runs AI workloads across 20+ clouds, Kubernetes, Slurm, and on-prem — deciding where compute should live, keeping jobs and clusters consistent across unreliable infrastructure, and recovering automatically when instances disappear. We're looking for an engineer to own this core end to end, set its technical direction and build new features to make it more robust, performant, and user-friendly. The core you own is what lets SkyPilot run the largest AI workloads in the world.
What you'll do
- Design and build the future of AI systems: you will solve some of the hardest problems in distributed AI systems to make SkyPilot the standard solution running frontier AI workloads.
- Shape the technical direction of SkyPilot: the modules, interfaces, and invariants the rest of engineering builds on — and raise the bar for how we design and operate the system.
- Build enhancements and new components to evolve SkyPilot with better support of a wide range of AI and batch workloads.
- Engage with users: Opportunity to work closely with our users and customers to make their use cases successful; to grow our open-source community; to gain visibility for your work via public tutorials, blog posts, and/or talks.
What we're looking for
- You've designed, built, and operated distributed systems at production scale.
- Deep systems fundamentals: scheduling, state machines, coordination, and fault tolerance.
- Python/Go fluency and strong reasoning skills about concurrent, distributed code.
- You reach for the simplest design that survives contact with real workloads — and can explain why.
- Experience with cloud provider control planes (AWS / GCP / Azure), Kubernetes internals, or operating across multiple clouds, and open-source or AI infrastructure contributions (e.g. KAI, KubeRay, Kueue, KServe).
What we offer
- Competitive compensation and equity
- Comprehensive medical, dental, vision coverage for you and your dependents
- The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
- A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
- Gourmet lunch & dinner for the team to do their best work
Location: San Mateo, CA. Remote will be considered for exceptional candidates.