Build and optimize distributed training infrastructure for pre-training and large-scale RL workloads using the prime-rl framework. Design low-level performance optimizations and improve end-to-end efficiency across compute, memory, and networking layers.
Prime Intellect
15 Remote Job Openings at Prime Intellect
Research Engineer - Reinforcement Learning
Prime Intellect
·
Full Time
·
23 days ago
Prime Intellect
Lead research in building massive-scale synthetic data generation pipelines and optimizing AI inference workloads for performance and cost. Contribute to open-source RL frameworks and publish findings in top-tier AI conferences like ICML and NeurIPS.
Open Application for Unconventional Talent
Prime Intellect
·
Full Time
·
a month ago
Prime Intellect
Build the open superintelligence stack and infrastructure for frontier AI labs. Develop open-source models and a full-stack platform for post-training, agent workflows, and autonomous research.
Lead the growth organization by overseeing sales, marketing, partnerships, and customer success to connect technology to the market. Define GTM strategies for RL infrastructure and close large-scale compute and post-training contracts.
Design and implement next-generation AI agents and post-training methods like RLHF and GRPO to align models with real-world workloads. Act as a technical bridge between customers and research teams to translate applied data into product and research priorities.
Collaborate with experts to develop state-of-the-art open language models, coding agents, and scientific discovery models. Contribute to the Prime Intellect platform to democratize AI and deploy large-scale models across thousands of GPUs.
Build and maintain the company's data warehouse and pipelines to create a live, accurate picture of compute supply and demand. Develop data models and dashboards to enable cross-functional teams to track utilization and capacity planning.
Technical Account Manager - AI Infrastructure
Prime Intellect
·
Full Time
·
3 months ago
Prime Intellect
Manage a portfolio of enterprise customers to ensure successful adoption, retention, and expansion of AI infrastructure services. Act as a technical partner to optimize training and inference workloads while translating customer needs into product feedback.
Member of Technical Staff - Full Stack Software Engineer
Prime Intellect
·
Full Time
·
3 months ago
Prime Intellect
Build and own the developer-facing platform, APIs, and web interfaces for AI workload management. Develop backend services in Python and create real-time monitoring tools for training and deploying frontier models.
Member of Technical Staff - Training Platform
Prime Intellect
·
Full Time
·
3 months ago
Prime Intellect
Design and operate Kubernetes-based training and inference orchestration across multi-cloud GPU fleets. Build developer-facing surfaces for job submission, monitoring, and model management using a modern full-stack.
You will own the analytical foundation for global compute markets, including pricing supply, modeling economics, and evaluating hardware generations. You will partner with leadership to drive strategic capital allocation and manage commercial diligence with cloud providers.
You will own the end-to-end compute strategy, including sourcing, economics, and contracting for GPU capacity to power the company's AI infrastructure. You will also partner with research and engineering teams to align supply with training roadmaps and technical requirements.
The role involves owning the security posture for the company's AI infrastructure, including threat modeling, secure architecture, and incident response. You will work directly with engineering and research teams to embed security into the stack and manage external security audits.
You will define the brand narrative and lead developer relations to drive adoption of Prime Intellect's open superintelligence infrastructure. This involves managing content engines, executing product launches, and building a high-performing marketing team.
The role involves building and optimizing the systems infrastructure for large-scale Reinforcement Learning and distributed training workloads, focusing on improving end-to-end efficiency across compute, memory, networking, and scheduling layers. Responsibilities include designing low-level performance optimizations like kernels and communication paths, and shaping the architecture of the RL training stack.