Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
You will maintain and scale Proxmox-based GPU cloud infrastructure while building and managing APIs to support customer workloads. The role involves diagnosing and resolving complex technical issues, automating recurring tasks, and ensuring the reliability of the control plane.
Massed Compute is building a modern GPU cloud platform for AI and high-performance compute workloads. Customers choose us when they need high-performance infrastructure with more flexibility, visibility, and support than they get from traditional options. Our work sits at the intersection of AI infrastructure, product design, systems thinking, and customer obsession.
We run a multi-tenant GPU cloud: physical servers in several data centers, a virtualization layer on top of them, and the APIs customers use to rent capacity. This role covers both halves — you'll work in a hypervisor and a switch config one day and in a backend service and a SQL query the next.
Work is ticket-driven and AI-assisted. We expect engineers to use agentic coding tools as a normal part of the job, and to write clearly enough that the next person can follow the reasoning. Customer-reported problems land here too, so some of the work is direct: reproduce it, fix it, tell the customer what happened.
Run and maintain Proxmox clusters across multiple sites — nodes, networking, hardware faults
Build and maintain APIs and MCPs, including the public customer-facing surface
Keep the control plane and the actual infrastructure in agreement; automate away recurring issues
Investigate and resolve customer-reported issues, and communicate the resolution back
Fix multi-tenancy, security, and billing correctness issues as they surface
Work tickets end to end: diagnose, fix, verify, write it up
Evaluate vendors and architecture options, and document the decision
API development and maintenance
SQL — querying, schema work, and reasoning about transactions and concurrency
Strong Linux and networking fundamentals
Backend development in at least one server-side language
Hands-on experience using agentic coding tools — Claude Code, Cursor, Codex, or similar — as part of your normal workflow
Practical experience using MCP servers, and a working understanding of how they fit into an agentic workflow
Customer-facing troubleshooting — you can work a problem with the person who reported it and explain the outcome without jargon
Comfort with a ticket-driven workflow and clear written communication
3+ years in infrastructure, platform, or SRE-adjacent engineering
Proxmox (or comparable virtualization) administration, including its API
Building MCP servers
Experience in Node.js/JavaScript and Python
GPU infrastructure: NVIDIA drivers, device passthrough, hugepages, firmware
Data center hands-on work: cabling, switch configuration, procurement
Billing or metering systems
Object storage and FUSE performance work
OAuth 2.0, OpenTelemetry
On-call and incident command experience
Recurring alert classes shrink instead of holding steady
Infrastructure changes land cleanly, without customer-visible impact
Provisioning, teardown, and billing paths tell the truth about what they did
Fewer surprises reach customers before they reach us
Equal Opportunity
Massed Compute is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, veteran status, or any other characteristic protected by law.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Platform Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”