Own the GPU and ML-systems layer, focusing on accelerator scheduling, utilization, and health monitoring across clouds and Kubernetes. Build optimizations for large-scale pre-training, high-throughput inference, and deepen integrations with frameworks like vLLM and PyTorch.
About SkyPilot
SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."
SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.
The role
SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads. A few points of GPU utilization here can save a team millions in compute and days on every training run.
What you'll do
- Own GPU scheduling, utilization and health: how SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery.
- Build optimizations for training and serving: Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals.
- Make the AI stack run great out of the box: deepen integrations with vLLM, PyTorch, Slime, and the frameworks teams use for pre-training and high-throughput inference.
What we're looking for
- Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure.
- Strongly preferred: Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe).
- You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after.
- Strong Python, and comfort reaching into systems-level and GPU-adjacent details.
- You care about squeezing most from the compute available to you
- Experience operating large-scale training or high-throughput inference in production
What we offer
- Competitive compensation and equity
- Comprehensive medical, dental, vision coverage for you and your dependents
- The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
- A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
- Gourmet lunch & dinner for the team to do their best work
Location: San Mateo, CA. Remote will be considered for exceptional candidates.