Member of the Technical Staff - Platform

 Posted 4 days ago
     
2-5 years experience
Apply Now

Please mention DailyRemote when applying

AI Summary

You will build and operate the control plane for a large-scale GPU fleet, managing the automated lifecycle of clusters from bare metal to customer-ready. This includes maintaining Kubernetes and Postgres systems while scaling infrastructure to support thousands of nodes.

Member of the Technical Staff, Platform

Location: North America Remote / San Francisco, CA · Full-Time

About Andromeda

Compute is the most sought-after resource in the world, yet it still trades like commercial real estate: year-long contracts, manual fulfillment, capacity sitting idle because nobody can move it. We are building the liquidity layer at Andromeda.

Andromeda was founded by Nat Friedman and Daniel Gross to give startups the scaled AI infrastructure once reserved for hyperscalers. The first cluster filled almost instantly. The years since went into the platform that makes compute liquid: it deploys into foreign datacenters and turns the hardware it finds into clusters that leading AI labs train on.

Today Andromeda operates compute for 80+ customers across 30+ capacity providers, with tens of thousands of GPUs under management and billions of GPU-hours supported, on everything from A100 to GB300.

The Role

We are looking for engineers to build and operate the control plane that runs our fleet. Your responsibilities will include:

  • The automated systems that take a cluster from bare machines to customer-ready

  • Machine lifecycle between tenants: join, wipe, verify, rejoin

  • Operating Kubernetes and Postgres across the fleet

  • Contributing to our custom Kubernetes operators

  • Scaling clusters from tens of nodes to thousands

  • Participating in on-call rotations

Requirements

  • Impressive technical work you can go deep on, with impact in the world. That can take three years or twenty.

  • 2+ years of on-call experience for critical production services

  • Deep Kubernetes experience

  • Strong Linux fundamentals: kernel, cgroups, containers, networking, storage

  • Experience operating databases, monitoring, CI/CD, and cloud infrastructure at scale

  • Familiarity with fleet management and capacity planning

Prior experience writing operators, or with GPUs, HPC scheduling, or bare-metal hardware is nice to have, but not required.

Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Similar Jobs

See all Remote Others jobs →

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Others

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified