For Employers

SkyPilot

Member of Technical Staff, GPU / ML Systems

Posted 2 months ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

Own the GPU and ML-systems layer, focusing on accelerator scheduling, utilization, and health monitoring across clouds and Kubernetes. Build optimizations for large-scale pre-training, high-throughput inference, and deepen integrations with frameworks like vLLM and PyTorch.

About SkyPilot


SkyPilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so SkyPilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."

SkyPilot (10k+ GitHub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.

The role


SkyPilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make SkyPilot the fastest, most cost-efficient place to run demanding AI workloads. A few points of GPU utilization here can save a team millions in compute and days on every training run.

What you'll do


  • Own GPU scheduling, utilization and health: how SkyPilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery.
  • Build optimizations for training and serving: Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals.
  • Make the AI stack run great out of the box: deepen integrations with vLLM, PyTorch, Slime,  and the frameworks teams use for pre-training and high-throughput inference.

What we're looking for


  • Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure.
  • Strongly preferred: Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe).
  • You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after.
  • Strong Python, and comfort reaching into systems-level and GPU-adjacent details.
  • You care about squeezing most from the compute available to you
  • Experience operating large-scale training or high-throughput inference in production

What we offer


  • Competitive compensation and equity
  • Comprehensive medical, dental, vision coverage for you and your dependents
  • The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
  • A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
  • Gourmet lunch & dinner for the team to do their best work

Location: San Mateo, CA. Remote will be considered for exceptional candidates.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Director of SEO & AI Search (Remote - West Coast Only & Pacific Time Hours)

Full Time United States Software Development

Senior AWS Cloud Engineer (Remote, Full-Time) [AS314]

Full Time India Software Development

Cloud Engineer - Frontend-leaning (Remote, Full-Time) [AS315]

Full Time India Software Development

Machine Learning Software Engineer / Senior Machine Learning Software Engineer

Full Time United Kingdom Software Development

Senior Inpatient Facility Certified Medical Coder

Full Time United States $23.89 - $42.69 per hour Software Development

Principal Software Development Engineer, Backend

Full Time United States $181K - $305K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 217,253+ Jobs in Software Development

Answer easy questions

Answer easy questions

217,253+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified