Architect and build SkyPilot's commercial managed control plane, focusing on the API server, multi-tenancy, and enterprise-grade platform services. Ensure the reliability, scalability, and security of the hosted infrastructure to support large enterprise GPU fleets.
About SkyPilot
SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."
SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.
The role
SkyPilot Platform is a managed control plane on top of the open-source core — a hosted API server that gives infrastructure teams unified scheduling, governance, and security across all their clusters and clouds, while compute stays in their own environment (BYOC). We're looking for an engineer to own this control plane end to end — the API server, multi-tenancy, and the services that turn SkyPilot into a product teams run their infrastructure on. This is the layer that makes a large enterprise comfortable putting its whole GPU fleet on SkyPilot.
What you'll do
- Architect SkyPilot's commercial platform from the ground up: the managed API server (hosted control plane), control-plane / data-plane separation, and the multi-tenant foundation teams run their compute on.
- Scale and operate the control plane: keep the hosted API server reliable and highly available — monitoring, alerting, seamless upgrades, and the observability teams depend on.
- Build enterprise-grade platform services: SSO, RBAC, secrets management, usage accounting, and SOC 2-grade security.
What we're looking for
- You've built SaaS / cloud platforms from zero to one, and you're equally comfortable in the cloud and Kubernetes infrastructure beneath them.
- Deep hands-on experience with major clouds' compute / networking / IAM APIs and Kubernetes internals (the API, controllers/operators, scheduling).
- Hands-on with core platform services: SSO, authentication and RBAC, API gateway, usage metering and billing, CI/CD.
- Fluent in a cloud-native stack — e.g. Go, Kubernetes, gRPC, PostgreSQL, Terraform — and comfortable in Python.
- You care about reliability, security, and simplicity in equal measure.
- Experience taking an open-source project into a commercial product, multi-cloud / hybrid systems, or GPU provisioning across clouds.
What we offer
- Competitive compensation and equity
- Comprehensive medical, dental, vision coverage for you and your dependents
- The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
- A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
- Gourmet lunch & dinner for the team to do their best work
Location: San Mateo, CA. Remote will be considered for exceptional candidates.