For Employers

Inferact

Member of Technical Staff, Cloud Orchestration (Remote)

Posted an hour ago
$200K - $400K per year
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

You will design and maintain the operational backbone for vLLM, including cluster management, deployment automation, and production monitoring. Your work will ensure that AI model inference systems are observable, debuggable, and highly reliable at a massive scale.

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.

  • Strong experience with Kubernetes and container orchestration at scale.

  • Experience designing and implementing custom Kubernetes operators.

  • Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).

  • Experience managing GPU clusters and debugging hardware issues.

  • Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.

Preferred qualifications:

  • Experience with ML-specific orchestration tools (Ray, Slurm).

  • Knowledge of GPU scheduling, multi-tenancy, and resource optimization.

  • Familiarity with vLLM deployment patterns and configuration.

  • Track record of improving operational reliability for ML systems.

Bonus points if you have:

  • Experience deploying inference systems on large-scale GPU (1,000+) clusters.

Logistics

  • Location: Fully remote, worldwide. We're timezone-flexible but expect regular overlap with Pacific Time for critical syncs.

  • Compensation: We offer competitive compensations (salary + equity) compared to the local market conditions.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Architect - Hybrid Cloud

Full Time Philippines Software Development

Senior Engineer (Datacenter)

Full Time Philippines Software Development

MDR Security Engineer (India)

Full Time India Software Development

Oracle GTM Consultant

Full Time Worldwide Software Development

Backend Developer (.NET)

Full Time Croatia Software Development

Global Reporting & Analytics Analyst

Full Time India Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Software Development

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified