Member of Technical Staff, GPU / ML Systems

 Posted 11 hours ago
     
5-10 years experience
Apply Now

Please mention DailyRemote when applying

AI Summary

Own the GPU and ML-systems layer, focusing on accelerator scheduling, utilization, and health monitoring across clouds and Kubernetes. Build optimizations for large-scale pre-training, high-throughput inference, and deepen integrations with frameworks like vLLM and PyTorch.

About SkyPilot


SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."

SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.

The role


SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads. A few points of GPU utilization here can save a team millions in compute and days on every training run.

What you'll do


  • Own GPU scheduling, utilization and health: how SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery.
  • Build optimizations for training and serving: Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals.
  • Make the AI stack run great out of the box: deepen integrations with vLLM, PyTorch, Slime,  and the frameworks teams use for pre-training and high-throughput inference.

What we're looking for


  • Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure.
  • Strongly preferred: Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe).
  • You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after.
  • Strong Python, and comfort reaching into systems-level and GPU-adjacent details.
  • You care about squeezing most from the compute available to you
  • Experience operating large-scale training or high-throughput inference in production

What we offer


  • Competitive compensation and equity
  • Comprehensive medical, dental, vision coverage for you and your dependents
  • The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
  • A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
  • Gourmet lunch & dinner for the team to do their best work

Location: San Mateo, CA. Remote will be considered for exceptional candidates.

Similar Jobs

See all Remote Software Development jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Software Development

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified