You will embed with customer teams to onboard, tune, and debug large-scale AI training and inference workloads on GPU clusters. You are also responsible for building automation and infrastructure to ensure the reliability and performance of these distributed systems.
Andromeda Cluster
13 Remote Job Openings at Andromeda Cluster
Member of the Technical Staff - Product
Andromeda Cluster
·
Full Time
·
4 days ago
Andromeda Cluster
You will own the backend systems responsible for capacity quoting, reservation, funding, and billing across the marketplace. Additionally, you will manage the full API surface and infrastructure operations, including metering and customer-facing dashboards.
Member of the Technical Staff - Platform
Andromeda Cluster
·
Full Time
·
4 days ago
Andromeda Cluster
You will build and operate the control plane for a large-scale GPU fleet, managing the automated lifecycle of clusters from bare metal to customer-ready. This includes maintaining Kubernetes and Postgres systems while scaling infrastructure to support thousands of nodes.
Member of the Technical Staff - Systems
Andromeda Cluster
·
Full Time
·
4 days ago
Andromeda Cluster
You will design and build high-performance systems across storage, networking, virtualization, and container runtimes. Additionally, you will participate in on-call rotations to respond to production incidents.
The role involves building and owning the revenue operating system, including GTM tech stack optimization, forecasting infrastructure, and deal desk processes. You will act as the cross-functional glue between sales, supply, legal, and finance to remove friction and scale revenue operations.
Technical Program Manager - Provider Management
Andromeda Cluster
·
Full Time
·
13 days ago
Andromeda Cluster
Own provider relationships and execution, managing onboarding, capacity rollouts, and hardware quality issues. Act as the incident commander for major provider-side failures, coordinating responses between SREs, providers, and customers.
You will lead financial planning, performance management, and deal execution for the company's global GPU compute portfolio. This includes building financial architecture, managing lender relationships, and overseeing capital allocation for infrastructure projects.
You will build and own the end-to-end partnership lifecycle, including sourcing, negotiating, and managing relationships with GPU providers, AI labs, and VC funds. You will also drive cross-functional coordination to ensure partnership agreements are successfully integrated into the company's infrastructure and sales operations.
Senior Site Reliability Engineer - AI Infrastructure
Andromeda Cluster
·
Full Time
·
4 months ago
Andromeda Cluster
You will design, operate, and debug large-scale GPU infrastructure for distributed AI training and inference. You will also serve as the primary technical partner for customers, ensuring the reliability and performance of high-speed interconnects and compute clusters.
Member of the Business Staff - Compute Markets
Andromeda Cluster
·
Full Time
·
4 months ago
Andromeda Cluster
The role involves matching incoming sales leads with internal and external compute capacity to maximize resource utilization across the platform. This includes sourcing and onboarding new global compute suppliers based on customer needs and market trends.
The Infrastructure Manager will be responsible for matching incoming sales leads with internal and external compute capacity while maximizing the utilization of existing compute resources. This role involves sourcing and onboarding new global compute suppliers and developing proactive compute strategies based on market intelligence.
Software Engineer - AI Infrastructure
Andromeda Cluster
·
Full Time
·
5 months ago
Andromeda Cluster
The role involves designing and developing core platform components, focusing on infrastructure orchestration, provisioning, and lifecycle management solutions across diverse infrastructure types. Responsibilities also include translating customer usage into product requirements and enhancing platform reliability and performance.
General Interest - Experience w/ AI Infrastructure
Andromeda Cluster
·
Full Time
·
5 months ago
Andromeda Cluster
This is a general interest posting for individuals with firsthand experience in AI infrastructure who do not see a currently open role matching their qualifications. The company will review submitted resumes for future opportunities that align with the candidate's background.