You will define the go-to-market strategy and solution narratives for AI-driven retail and e-commerce workloads on the Nebius AI Cloud. You will also create sales enablement assets and collaborate with cross-functional teams to drive customer adoption and revenue growth.
Nebius
138 Remote Job Openings at Nebius
You will define and own reliability goals for network services while building automation to improve site readiness and inter-site connectivity. Additionally, you will lead incident response efforts and design safer change workflows to ensure high availability across the network infrastructure.
Lead and develop a team of Forward Deployed Engineers to build high-quality integrations, reference architectures, and proofs of concept for AI workloads. Act as a technical escalation point and bridge the gap between partner needs and internal product development.
You will coordinate customer-related activities across the full lifecycle, including discovery, onboarding, and production support. You will lead cross-functional initiatives to ensure technical teams work as a unified system to deliver results and improve processes.
Build and maintain services and tooling that automate the network lifecycle, including provisioning, drift detection, and operational verification. Develop observability systems and integrate network data to ensure safe, scalable, and transparent network operations.
Lead and develop a high-performing strategic sales organization to secure major AI infrastructure agreements with global accounts. Act as a strategic advisor to executive leadership to shape long-term commercial strategy, ecosystem partnerships, and market growth initiatives.
You will lead the mechanical engineering evaluation of prospective data center sites, including cooling capacity analysis and equipment feasibility reviews. You will also partner with cross-functional teams to define mechanical strategies and ensure technical readiness before construction begins.
The IT Risk & Controls Manager will act as an embedded partner to engineering teams to design and implement scalable IT SOX controls across the cloud platform. You will lead audit readiness, evaluate control gaps, and collaborate with stakeholders to ensure compliance while supporting operational efficiency.
You will lead electrical due diligence, utility interconnection analysis, and equipment feasibility reviews to inform strategic infrastructure acquisition decisions. Additionally, you will coordinate with cross-functional teams to define power distribution topologies and ensure construction readiness for AI cloud infrastructure.
The MEP Engineer will support technical due diligence and feasibility studies for AI infrastructure projects, including cooling and power distribution evaluations. They will also coordinate with cross-functional teams and vendors to ensure technical readiness and documentation for new data center sites.
The Network Planning Technical Project Manager will own the connectivity lifecycle from site assessment to infrastructure delivery, acting as the primary bridge between internal data center teams and external telecommunications partners. Responsibilities include managing fiber pathway architecture, negotiating commercial agreements with ISPs, and overseeing geospatial documentation for global network integration.
You will manage complex technical programs across hardware and software engineering teams to scale the Nebius compute platform. This involves identifying dependencies, mitigating risks, and establishing processes to support the deployment of large-scale GPU clusters.
You will manage the connectivity lifecycle from site assessment to infrastructure delivery, acting as the bridge between internal data center teams and external telecommunications partners. Responsibilities include leading fiber pathway architecture design, negotiating commercial agreements, and maintaining geospatial documentation for network routes.
You will build a comprehensive view of self-service customers to identify growth opportunities and translate commercial plans into capacity forecasts. Additionally, you will develop pricing models and partner with cross-functional teams to drive product and operational improvements.
The Senior Technical Project Manager will coordinate complex cross-functional projects, including new region launches and the end-to-end delivery pipeline for AI models. They will manage dependencies, risks, and engineering execution to ensure ambitious technical goals are met predictably.
Drive the adoption and consumption of the Token Factory AI inference platform by managing the full sales cycle from prospecting to close. Partner with Sales Engineering to convert proofs-of-concept into production commitments and scale consumption within AI-native and enterprise accounts.
Lead the development of AI infrastructure projects from executed agreements through construction readiness, managing roadmaps, budgets, and risk registers. Coordinate with utilities, municipalities, and cross-functional teams to ensure projects are fully permitted and ready for delivery.
Lead multidisciplinary technical due diligence for powered land and build-to-suit AI infrastructure opportunities to identify technical risks before capital commitment. Establish repeatable standards, playbooks, and consultant networks to scale the infrastructure acquisition process across the Americas.
The role involves analyzing and administering employee compensation programs to ensure competitive and equitable pay practices. Key duties include conducting market pricing, maintaining salary structures, and managing merit and bonus planning cycles.
The role involves analyzing and administering employee compensation programs to ensure competitive and equitable pay practices. Key duties include conducting market pricing, maintaining salary structures, and managing merit and bonus planning cycles.
Drive the execution of high-priority strategic initiatives, including product launches and GTM programs, by aligning cross-functional teams. Develop operating mechanisms and frameworks to translate ambiguous goals into structured execution plans for senior leadership.
Drive complex initiatives for the Virtual Private Cloud team, including launching VPC in new regions and improving platform scalability. Coordinate delivery across multiple engineering teams by turning technical objectives into executable programs and managing dependencies.
Manage the end-to-end deal lifecycle and partner with account executives to structure complex enterprise contracts. Own the GTM tech stack administration and serve as the primary liaison between Sales, Finance, and Legal.
Own server and rack-level production projects for large-scale AI datacenter infrastructure, connecting capacity planning with manufacturing and deployment. Manage the end-to-end lifecycle of hardware procurement, production tracking with ODM partners, and risk mitigation to ensure timely delivery.
The role focuses on optimizing commercial operations processes across contracting, billing, and sales support to ensure data integrity. Responsibilities include validating commercial requests, managing the quote-to-signature lifecycle, and coordinating with Finance to resolve invoicing discrepancies.
Lead the end-to-end design process for new data center builds and expansions, focusing on high-density GPU clusters. Coordinate with internal teams and external partners to ensure designs optimize power, cooling, and scalability while meeting regulatory standards.
Build and maintain GTM dashboards and SQL queries to analyze pipeline, bookings, and productivity trends. Partner with leadership to turn complex data into actionable insights and practical recommendations for business growth.
Provide high-level technical support for AI engineers and architects by diagnosing issues across Tavily products and LLM applications. Collaborate with engineering and product teams to resolve bugs and improve internal diagnostics and documentation.
Develop fault-tolerant and reliable cloud services including managed databases, Kubernetes, and billing systems. Focus on building a hyperscaling platform for the global AI economy.
Serve as the technical interface between Nebius and strategic technology partners to design and develop integrated AI solutions. Responsibilities include building reference implementations, enabling partner engineering teams, and influencing the product roadmap based on partner feedback.
Build and test LLM-based solutions and applications using Token Factory's inference services. Assist with prompt engineering, model selection, and performance experiments to support proof-of-concept work.
Own the product roadmap and delivery for Dedicated Endpoints for AI Models Inference. Partner with engineering and research to design high-performance tools and drive product-market fit through user research and customer onboarding.
Lead the creation and operation of a Vulnerability Operation Center, managing the full lifecycle from detection to remediation across cloud and hardware stacks. Define processes, select tooling, and coordinate responses to critical zero-day vulnerabilities while driving accountability across engineering teams.
The Senior HRBP will partner with USA regional leadership to build organizational capability and translate growth plans into practical workforce and hiring strategies. They will lead organizational design, coach leaders on performance and talent reviews, and manage people risks to ensure effective scaling.
Build and lead the internal red team function to conduct adversarial simulations and penetration testing across the Nebius cloud platform. Research novel attacks against GPU infrastructure and AI managed services to improve overall security posture.
Lead the Go-To-Market motion for Physical AI across EMEA by generating a qualified pipeline and managing strategic executive-level accounts. Develop repeatable sales playbooks and enable the regional field team to accelerate deal closure and drive compute consumption.
Lead complex, cross-functional GTM systems initiatives including CRM and marketing automation from discovery through implementation. Partner with Sales, Marketing, and Engineering to design scalable operating processes and governance models that drive growth.
Develop and optimize an inference and fine-tuning platform for various foundation models at massive scale. Focus on enhancing fine-tuning methodologies, identifying inference bottlenecks, and investigating low-precision training for modern hardware.
Manage complex, multi-stakeholder enterprise sales cycles to drive adoption of NVIDIA-based GPU infrastructure across priority verticals. Build executive relationships with C-level leaders and collaborate with strategic partners like NVIDIA and OEMs to execute co-sell motions.
Lead technical engagement and end-to-end solution ownership for strategic AI/ML customers across the Gulf region. Act as a trusted advisor to executive stakeholders to drive the adoption of Nebius AI Cloud through complex PoCs and architecture design.
Lead engagement and growth with strategic Managed Service Providers (MSP) and Value-Added Resellers (VAR) to position Nebius as a preferred partner. Develop partner enablement programs and collaborate with leadership to drive incremental revenue growth for GPU IaaS solutions.
AI Solution Architect - Educational Content Author, Nebius Academy
Nebius
·
Full Time
·
25 days ago
Nebius
Create technical educational content, including tutorials, sample code, and video sessions, to drive adoption of Nebius cloud infrastructure. Collaborate with academic partners to design and document solution architectures and provide technical expertise on GPU cloud technologies.
Develop technical courses on the Nebius AI Cloud for B2B and B2C audiences by translating SME materials into learner-friendly content. Collaborate with production teams to create visual aids and continuously improve courses based on student feedback and audits.
Ensure fault-tolerance, scale, and uninterrupted operations for hardware automation services. Implement and improve CI/CD processes while providing real-time visibility and control over data center infrastructure.
Own the end-to-end technical execution of infrastructure design and production rollout for strategic physical AI customers and ISV partners. Translate customer pain points into reusable platform capabilities to influence the core Nebius Physical AI roadmap.
Build and lead the Detection and Response capability from the ground up, owning detection engineering, threat intelligence, and incident response. The role involves architecting coverage across cloud and bare-metal environments and managing a growing team of analysts.
Own the vision, roadmap, and priorities for wide-area and underlay networking products, including data center fabric and WAN transport. Coordinate cross-company initiatives to define requirements for connectivity products and ensure high standards for throughput, latency, and reliability.
Own the day-to-day commercial relationships with strategic technology partners and named accounts within the Media and Entertainment vertical. Drive the expansion of the partner ecosystem by developing commercial frameworks, joint value propositions, and operational cadences.
Design and implement cutting-edge cloud infrastructure and MLOps solutions to help clients optimize AI pipelines. Act as a trusted technical advisor, conducting PoCs and workshops to educate clients on GPU cloud technologies.
Own the vision, roadmap, and product backlog for cloud storage services including block, file, and object storage. Coordinate cross-company initiatives to optimize data durability, performance, and cost for AI/ML and HPC workloads.
Act as the senior technical authority for customers using the Token Factory platform to design and implement optimized LLM inference and fine-tuning workflows. Partner with product and research teams to shape the platform roadmap and mentor other Solutions Architects.
Build and activate local developer ecosystems in North America by organizing builder-first events and managing ambassador programs. Act as the regional face of Nebius to engage AI builders and provide feedback to product and marketing teams.
Own the operational health and readiness of COLO and BTS data center sites, ensuring SLAs are met and maintenance is executed. Act as the primary interface between the company and data center landlords to translate site operations into reliable infrastructure.
The PR Manager will lead the media relations strategy, managing relationships with journalists and industry analysts to build brand reputation. Responsibilities include writing press releases, coordinating executive visibility, and partnering with Product Marketing for launch PR.
Own the full enterprise sales cycle from prospecting to closing for AI-powered search and data capabilities. Partner with technical and executive stakeholders to scale Tavily's presence within large organizations.
The Product Marketing Manager will own the positioning, messaging, and go-to-market strategy to ensure market value is understood. They will create sales enablement materials and conduct competitive research to support product growth.
Manage the full post-sale lifecycle for commercial accounts, focusing on onboarding, adoption, renewal, and expansion. Act as the voice of the customer to provide actionable insights to Product and Go-to-Market teams.
Own the end-to-end product strategy, roadmap, and delivery for a specific slice of the AI Compute Platform. Drive cross-functional execution across engineering and SRE teams to deliver hyperscaler-quality APIs and infrastructure outcomes.
Build and scale core services for a unified observability ecosystem covering logs, metrics, traces, and alerting. Focus on high-volume telemetry ingestion, distributed storage, and AI-assisted troubleshooting to support a global AI cloud platform.
Serve as a technical advisor helping customers design, deploy, and scale AI solutions and large-scale GPU workloads. Act as a bridge between customers and product teams to resolve complex AI/ML issues and drive product growth.
Design, build, and maintain automation workflows and integrations across IT systems and SaaS platforms. Automate operational tasks such as user management and data synchronization while ensuring secure implementation and documentation.
Establish and lead the customer success function for EMEA, guiding enterprise customers through onboarding, implementation, and long-term adoption of the AI search platform. Act as a strategic partner to executive sponsors and a bridge between customers and internal product teams to influence the roadmap.
Design and implement embedded firmware for server management, telemetry, and control systems for GPU and HPC platforms. Maintain custom OpenBMC firmware and collaborate with hardware engineers to validate and optimize low-level drivers.
The role involves building and maintaining ASPM tools to identify and remediate application security vulnerabilities. You will collaborate with development teams to integrate security best practices into the SDLC and conduct penetration testing.
Identify and secure optimal locations for data center developments and colocation expansions through market analysis and technical due diligence. Manage the full colocation lifecycle, including vendor RFPs, contract negotiations, and performance monitoring against SLAs.
Develop the control plane and lifecycle automation for Managed PostgreSQL while tuning database internals for AI workloads. Build migration tooling and drive the integration of vector search capabilities within the Nebius AI Cloud stack.
Maintain and grow DevTools systems, including GitLab and TeamCity, to support large-scale AI cloud infrastructure. Focus on building fault-tolerant architectures and improving system performance based on user feedback.
Lead the technical product strategy and GTM motion for Retail & Commerce AI, serving as the principal technical partner for strategic lighthouse accounts. Translate customer needs into scalable platform requirements and architect solution patterns for the broader market.
Design and implement cloud infrastructure and MLOps solutions for clients, acting as a trusted technical advisor. Conduct PoCs, workshops, and optimize pipeline performance to ensure efficient utilization of GPU cloud resources.
Design and implement cloud infrastructure and MLOps solutions for clients, acting as a trusted technical advisor. Responsibilities include conducting PoCs, optimizing pipeline performance, and collaborating with product and marketing teams.
Build and scale Nebius' ISV and AI partner ecosystem from the ground up to drive net-new customer acquisition and revenue. Develop repeatable co-sell motions and lead joint go-to-market execution with strategic partners.
The Director of Product, Ecosystem is responsible for mapping and engaging key AI companies to build a robust partner ecosystem across various platform layers. This includes prototyping technical integrations and translating external market signals into product capabilities and strategic investment decisions.
Design and develop internal web applications to automate data center and GPU cloud infrastructure operations. Collaborate with infrastructure and platform teams to translate operational requirements into scalable frontend solutions.
Drive go-to-market execution and manage strategic client relationships within the Digital Health, Medical Devices, and Medical Imaging segments across EMEA. Collaborate with Account Executives to qualify opportunities and position AI/cloud solutions to meet healthcare regulatory and business needs.
Investigate and resolve complex technical issues involving Linux, Kubernetes, and GPU-based AI workloads in customer environments. Act as a senior escalation point for production incidents and collaborate with engineering to develop long-term fixes and automation tools.
Plan and implement cost-effective mechanical designs and distribution systems for AI data centers. Collaborate with internal partners and external vendors to validate installation and operational performance of mechanical systems.
Plan, design, and oversee the implementation of electrical engineering systems for AI data centers. Collaborate with internal technical teams and external vendors to ensure quality benchmarks are met on time and within budget.
Lead and support the benchmarking of GPU platforms to evaluate performance for machine learning and AI workloads. This includes profiling GPU performance at the system and kernel level and optimizing ML workloads to resolve bottlenecks.
Serve as the primary technical advisor for strategic GPU Cloud customers to design and scale AI solutions. Collaborate with sales and product teams to optimize GPU performance and align product features with customer requirements.
Serve as a technical advisor helping clients design, deploy, and scale AI solutions and manage large-scale GPU workloads. Collaborate with sales and product teams to resolve complex AI/ML issues and align product features with customer requirements.
Design and prototype integrations between partner products and the Nebius platform to define scalable reference architectures. Translate external integration findings into actionable product requirements to shape the platform roadmap.
Serve as a technical advisor helping strategic customers design, deploy, and scale large-scale GPU workloads for AI solutions. Collaborate with sales and product teams to drive growth and relay customer feedback for product enhancement.
Lead the design, implementation, and optimization of global enterprise wired and wireless networks to ensure scalability and security. Manage remote access systems and drive automation-first operations for network monitoring and connectivity.
You will co-own the Serverless AI product roadmap, defining technical requirements and making strategic build-versus-buy decisions for infrastructure capabilities. Additionally, you will drive customer adoption through technical content, market analysis, and direct engagement with ML engineers.
Design and implement LLM-based solutions using the Nebius Token Factory inference platform to drive business value. Collaborate with product and engineering teams to shape the platform roadmap and guide customers from POC to production.
The Technical Account Manager will lead the transition of customer AI workloads from proof-of-concept to stable, scalable production environments. They will act as a trusted technical partner to optimize performance, cost-efficiency, and reliability while managing incidents and providing feedback to internal product teams.
You will tune and troubleshoot GPU clusters and InfiniBand networks to ensure optimal performance in high-performance computing environments. Additionally, you will integrate new hardware into the infrastructure and enhance automation systems for proactive monitoring and issue resolution.
You will analyze and optimize the performance of large-scale GPU clusters by identifying bottlenecks across hardware and software layers. Additionally, you will support performance-related escalations and contribute to hardware qualification and cluster validation.
The Technical Account Manager will lead the transition of customer AI workloads from proof-of-concept to production, ensuring stability, scalability, and cost-efficiency. They will act as a primary technical partner to resolve bottlenecks, manage incidents, and provide actionable feedback to internal product teams.
The Lead Systems HPC Engineer will analyze and optimize the performance of large-scale GPU clusters by identifying bottlenecks across hardware and software layers. They will also collaborate with infrastructure and vendor teams to integrate new hardware and support complex performance-related escalations.
You will drive the adoption of the Token Factory platform by creating demos, workshops, and technical guides for AI developers. Additionally, you will act as a bridge between the developer community and the internal product and engineering teams to gather feedback and improve the platform.
Drive planning and execution across multiple engineering and product teams to launch new external-facing cloud services. Coordinate service releases, manage stakeholder expectations, and align requirements between Sales, Partnerships, Legal, and Engineering.
You will initiate and drive cross-team projects while aligning conflicting priorities and stakeholder needs across engineering domains. Additionally, you will facilitate technical decision-making and establish scalable processes to support core cloud infrastructure.
You will manage large-scale projects for the Nebius Cloud Platform, coordinating across technical and business teams to scale infrastructure and support new hardware. This involves initiating cross-team projects, setting goals, facilitating decision-making, and aligning stakeholder needs up to the C-level.
The Enterprise Applications Engineer will manage, configure, and maintain the company's collaboration SaaS ecosystem, including Atlassian tools, Slack, and Zoom. They will also drive automation, ensure platform reliability, and partner with security teams to enforce governance and access controls.
You will be responsible for the administration, configuration, and continuous improvement of the Atlassian Cloud ecosystem, including Jira and Confluence. Additionally, you will collaborate with cross-functional teams to ensure business applications remain secure, efficient, and aligned with organizational standards.
You will be responsible for scaling distributed backend systems and optimizing AI infrastructure to ensure high performance and reliability. The role involves collaborating with cross-functional teams to build and maintain enterprise-grade AI integration platforms.
The lead will drive applied research across retrieval, ranking, and agent-centric search systems, focusing on designing and improving multi-stage retrieval pipelines and developing grounding approaches for LLMs using real-time web data. Responsibilities also include defining evaluation methodologies, leading experimentation on modern retrieval techniques, and working closely with engineering to deploy research into production at scale.
You will design, train, and deploy machine learning models for retrieval, reranking, and search relevance in production. This role involves working on systems operating at large scale and collaborating closely with engineering teams.
The Senior Network Engineer will be responsible for designing, building, and operating large-scale, high-performance data center networks supporting GPU-dense AI workloads, taking end-to-end ownership of service provider–grade and CLOS-based network infrastructure. Responsibilities include designing scalable architectures, owning deployment and lifecycle management of routing/switching infrastructure, and optimizing traffic engineering strategies.
The role involves spending roughly half the time in the field assisting new customers with POCs and technical onboarding, and the other half building prototypes, exploring emerging AI techniques, and translating field insights into product direction. Responsibilities include building demos across the portfolio, supporting customer validation, and feeding grounded feedback into the product roadmap.
The Senior ML Solutions Architect will design and implement LLM-based solutions utilizing Nebius Token Factory's inference services to achieve business value and support customer objectives. Responsibilities also include building production-ready applications with multimodal and domain-specific LLM APIs and collaborating with engineering teams to refine the platform based on client needs.
The manager will own the development and execution of the startup community engagement strategy across key markets, focusing on cultivating relationships with developers, accelerators, and founders. Responsibilities include designing scalable engagement programs, leading high-impact events, driving startup pipeline conversion, and developing enablement resources.
This role involves leading deep technical discovery with engineering teams to understand complex AI workload requirements, translating customer ambitions into production-feasible, scalable, and economically viable architectures. The engineer will partner closely with Sales to influence deal strategy, prevent misaligned commitments, and ensure high PoC-to-production conversion rates.
The Regional GTM Recruiting Manager will own the Go To Market recruiting strategy and execution across EMEA and APJ, leading and developing the regional recruiting team while partnering with commercial leadership to align hiring with revenue targets. Responsibilities include structuring workforce planning, supporting expansion into new regions, building scalable processes, and ensuring high-quality candidate experiences.
This role involves leading deep technical discovery with engineering teams to understand complex AI workload requirements, translating customer ambitions into scalable, production-feasible architectures, and identifying technical risks early. The Senior Sales Engineer will partner closely with Sales to influence deal strategy, ensure technical realism, and drive PoC-to-production conversion by defining measurable success criteria.
The Senior Delivery Deployment Engineer will own the end-to-end delivery, deployment, and production readiness of next-generation GPU platforms inside data centers, leading on-site rack bring-up and validating NVIDIA-based AI systems. Responsibilities include overseeing installation, troubleshooting complex hardware/OS/network issues, executing validation procedures, and coordinating on-site hardware repairs.
The role involves adapting YDB to fully utilize modern hardware like QLC NVMe drives, DPUs, and maximizing performance on standard devices like HDDs and TLC NVMe. Responsibilities also include reengineering YDB components using more efficient algorithms to address complex system challenges.
The Principal Solutions Architect will own end-to-end technical presales and solution delivery for strategic customers, acting as a trusted technical advisor to senior stakeholders on complex ML/AI infrastructure and MLOps architectures. Responsibilities also include driving advanced Proofs of Concept (PoCs), influencing the product roadmap, and mentoring other Solutions Architects.
The role involves serving as the technical interface between Nebius and strategic technology partners, designing and developing integrated solutions, building reference implementations, and enabling partner engineering teams. Responsibilities include owning the technical relationship with partners, architecting high-impact integration solutions, and translating feedback into product roadmap requirements.
The role involves building a distributed, fault-tolerant storage system for the hyperscaler cloud, focusing on high-performance block and filesystem storage capable of sub-millisecond latency and high throughput despite hardware failures. Key work areas include virtualization technologies like virtio-blk and QEMU, high-throughput networking over TCP and RoCEv2, and implementing storage system features like replication and self-healing.
As a Systems Administrator, you will maintain, troubleshoot, and integrate Linux-based systems that support production platforms. You will work closely with infrastructure, networking, and data center teams to ensure reliable services.
The role involves developing a data plane for Nebius Cloud network services and improving service performance and reliability. Additionally, the engineer will automate complex cross-service interaction scenarios and perform load testing for services.
As a Senior Backend Software Engineer, you will design, build, and operate backend services that are reliable, scalable, and performant. You will own services end-to-end, from architecture to production support.
As a Software Engineer, you will design and build systems for provisioning, configuring, testing, and managing physical hardware at scale. You will collaborate with hardware, networking, and data center operations teams to ensure robust and scalable platforms.
The Senior Software Engineer will design, build, and own backend systems for metrics and monitoring large-scale infrastructure. Responsibilities include evolving metrics pipelines and investigating production incidents.
The Data Center Technician Manager leads technical teams responsible for cabling, hardware installation, and data center operations. This role involves mentoring technicians and ensuring high-quality execution of installations and maintenance.
The Data Center Site Manager is responsible for ensuring the reliability, safety, capacity, and performance of a flagship U.S. site. This includes leading a multi-disciplinary operations team and managing various aspects of data center operations.
Ensure the reliability, availability, and performance of compute nodes running VMs. Troubleshoot complex production issues and lead incident response and root-cause analysis.
Design and develop a large-scale LLM training platform while maintaining optimal performance, scalability, and reliability of the ML infrastructure. Improve job scheduling strategies to minimize resource fragmentation.
Develop and optimize low-level kernels and runtime components for AI inference. Collaborate with ML and backend teams to optimize end-to-end execution.
Own key tracks in Cluster Experience, focusing on reliability, performance, and user experience for distributed ML workloads. Define product direction and drive cross-functional execution across various teams.
As a Cloud Solutions Architect, you will design and implement solutions for clients, providing technical expertise and guidance. You will also conduct workshops and help optimize pipeline performance for efficient cloud resource utilization.
The Infrastructure Security Engineer will design and implement security measures for cloud and on-premises infrastructure, identify and remediate vulnerabilities, and collaborate with teams to integrate security best practices. They will also stay updated on security threats and serve as a subject matter expert.
Design and implement LLM-based solutions using Nebius Token Factory’s inference services to drive business value. Collaborate with product and engineering teams to surface customer feedback and shape the platform roadmap.
The Senior HPC Cluster Engineer will tune the performance of GPU clusters and InfiniBand networks, analyze and troubleshoot issues, and integrate new hardware into the existing infrastructure. The role also involves enhancing automation systems for proactive monitoring and managing GPU devices and InfiniBand fabrics.
The Senior Hypervisor Engineer will optimize I/O for emulated devices, integrate the hypervisor with platform services, and improve guest system support. They will also advance open source virtualization and provide low-level security for the hypervisor.
The Senior Hypervisor Engineer will optimize I/O for emulated devices, integrate the hypervisor with platform services, and improve guest system support. The role also involves collaborating with the open-source community and providing low-level security for the hypervisor.
Ensure fault-tolerance, scale, and uninterrupted operations for the service. Use cutting-edge cloud technology to solve a variety of infrastructure problems.
Conduct experiments to efficiently train large language models and explore methods of guided generation and search. Mine relevant data at web scale and conduct experiments with various reinforcement learning configurations.
Technical Project Manager (Devtools and Observability Platform)
Nebius
·
Full Time
·
8 months ago
Nebius
The Senior Technical Project Manager will set up feedback loops from engineering teams, drive AI-related experiments, and enhance planning and processes. This hands-on role focuses on structuring and pushing initiatives forward rather than extensive reporting.
The role involves developing a data plane for Nebius Cloud network services and improving service performance and reliability. Additionally, the engineer will automate complex cross-service interaction scenarios and perform load testing for services.
Senior Site Reliability Engineer — AI Studio (Inference Platform)
Nebius
·
Full Time
·
8 months ago
Nebius
You will own the reliability, performance, and observability of the entire inference stack. This includes designing telemetry pipelines, tuning Kubernetes autoscalers, and creating automation for incident management.
The Compute Node team builds services for managing Virtual Machines on GPU servers and develops the Virtual Machine Scheduler for clusters with thousands of servers. This involves integrating with disk management and virtual networks across multiple data centers.
The Site Selection & Colocation Manager is responsible for identifying, evaluating, and securing optimal locations for new data center developments and colocation expansions. This includes conducting market assessments, coordinating technical due diligence, and managing vendor relationships.
The Site Selection & Colocation Manager is responsible for identifying, evaluating, and securing optimal locations for new data center developments and colocation expansions. This role involves collaboration with various teams to ensure that selected sites meet operational, technical, and business requirements.
Technical Project Manager (Devtools and Observability Platform)
Nebius
·
Full Time
·
9 months ago
Nebius
The Senior Technical Project Manager will set up feedback loops from engineering teams, drive AI-related experiments, and help mature planning and processes. This hands-on role focuses on structuring and pushing projects forward rather than extensive reporting.
The Talent Sourcer will be responsible for sourcing candidates for various roles within Data Centers, including technical positions. They will manage talent pools and collaborate with the recruitment team to ensure a positive candidate experience.