The manager will oversee international cloud compute supplier relationships, manage contract lifecycles, and run RFP processes for new capacity. They will also develop market intelligence and negotiation strategies to optimize costs and ensure supply chain efficiency for the global GPU fleet.
Together AI
13 Remote Job Openings at Together AI
Technical Support Engineer (Inference) - US Weekends
Together AI
·
Full Time
·
17 days ago
Together AI
You will act as a technical support engineer and customer-facing SRE to resolve complex issues with AI inference and fine-tuning services. You will also collaborate with engineering and product teams to improve system health, manage infrastructure changes, and drive the product roadmap.
Technical Support Engineer (GPU Clusters) - US Weekends
Together AI
·
Full Time
·
17 days ago
Together AI
Provide technical support for customers building AI solutions on Kubernetes GPU clusters while acting as a customer-facing SRE. Monitor cluster health, resolve complex infrastructure issues, and collaborate with engineering teams to improve product offerings.
Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking
Together AI
·
Full Time
·
17 days ago
Together AI
You will act as a technical partner to strategic customers, optimizing inference engines and guiding fine-tuning pipelines for production AI models. Additionally, you will influence the product roadmap by surfacing field insights and ensuring successful platform adoption.
Own and improve the end-to-end qualification process for new compute capacity to ensure it meets technical standards. Coordinate with engineering teams to validate hardware, networking, and operational resilience before recommending go/no-go decisions.
Maintain user-facing services and production systems using automation and sound engineering principles. Build and scale AI infrastructure using Ansible, Terraform, and Kubernetes while managing on-call rotations and monitoring systems.
Design and implement a scalable observability platform for metrics, logs, and traces to monitor system performance and GPU utilization. Develop automated monitoring, alerting systems, and custom infrastructure-as-code tools to enhance distributed tracing and incident response.
Provide analytical support for infrastructure strategy, focusing on capacity planning, compute sourcing, and site selection. Build dashboards and frameworks to monitor deployments and streamline operational workflows across Engineering and Finance.
Design and build infrastructure for model customization and evaluation, including backend services and job orchestration platforms. Collaborate cross-functionally to improve platform reliability, observability, and deployment tooling.
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AI
·
Full Time
·
3 months ago
Together AI
Design and deliver multi-petabyte storage systems and high-performance parallel filesystems optimized for AI training and inference workloads. Build Kubernetes-native storage operators and optimize end-to-end data paths to achieve high throughput and low latency.
Partner with AI research and engineering leadership to define and execute hiring strategies for world-class research teams. Lead full-cycle recruiting for specialized AI talent across academia and industry while refining technical interview processes.
Own the design, planning, and technical execution of whitespace environments for large-scale AI GPU clusters across a data center portfolio. Collaborate with MEP consultants and network engineers to optimize power, cooling, and physical layer requirements for high-density deployments.
You will optimize and fine-tune GPU code for better performance and scalability while collaborating with cross-functional teams to integrate GPU-accelerated solutions. Staying up-to-date with advancements in GPU programming techniques is also a key responsibility.