For Employers

Proxima

Principal ML Performance Engineer (GPU Optimization)

Posted a day ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

Profile and optimize training and inference for structural and generative models while scaling distributed training across large GPU clusters. Manage GPU cluster efficiency on GCP and develop benchmarking tools to support research team performance.

Principal ML Performance Engineer (GPU Optimization)

About Proxima

Proxima is a frontier AI and data generation company discovering the next generation of proximity therapeutics by making protein interactions programmable. Our platform brings together foundation-model machine learning, a scalable data generation engine, and a partnership track record exceeding $5B in collaborations across the world’s leading biopharma and tech organizations. We’ve recently closed an oversubscribed seed round with an elite group of VCs including DCVC, NVIDIA’s NVentures, AIX, Yosemite among others.

Neo-1 is our all-atom foundation model that combines state-of-the-art structure prediction and molecular generation in a single system. Neo-1 enables rapid exploration of chemical and structural space for high value, previously intractable targets, and in particular unlocks small molecule proximity therapeutics like molecular glues with AI for the first time.

In parallel, we are developing an advanced structural interactomics platform built on proprietary XLMS technology and a lab equipped with next-generation mass spectrometry instrumentation. This platform produces proteome-scale maps of protein interactions and helps identify small molecules that modulate proximity. Together with Neo-1, it creates an integrated system capable of co-folding protein complexes while generating candidate small molecules to influence those interactions.

Proximity-based therapeutics represent one of the most promising frontiers in modern drug discovery with the potential to treat previously intractable diseases and target ‘undruggable’ proteins. We’re building the tech and the team to make that happen. Come join us!

What you'll do

  • Profile and optimize training and inference for structural and generative models, including transformers, diffusion, and geometric deep learning

  • Write and tune custom kernels (CUDA, Triton) and use compilers (torch.compile, TensorRT, XLA) when beneficial

  • Scale distributed training across 32-64 nodes, employing FSDP, DeepSpeed, tensor and pipeline parallelism, and mixed precision

  • Reduce inference cost by optimizing memory scaling for large complexes, improving diffusion sampling efficiency, batching ragged inputs, and maximizing throughput across up to 1000 GPUs

  • Manage GPU cluster efficiency on GCP, focusing on scheduling, utilization, spot strategy, and cost reporting

  • Develop benchmarks and profiling tools for the research team

What we need

  • Minimum of 6+ years experience in ML systems, HPC, or performance engineering, with a BS/MS/PhD in CS, EE, or related field

  • Demonstrated ability to set technical direction beyond coding: selecting infrastructure, influencing research teams, and mentoring engineers

  • Deep knowledge of PyTorch internals with hands-on experience profiling and fixing real bottlenecks

  • Experience with CUDA and Triton, skilled at reading Nsight output, and strong understanding of memory bandwidth and occupancy

  • Experience with distributed training at multi-node scale

  • Strong proficiency in Python and C++

  • Able to name a model they made materially faster and quantify the improvement

Nice to haves

  • Experience in geometric deep learning, equivariant networks, or protein structure models such as AlphaFold, ESM, or RFdiffusion

  • Experience writing kernels for structure-model primitives, including triangle attention, triangle multiplicative updates, cuEquivariance, or FlashAttention for pair bias

  • Experience orchestrating large batch inference and managing Kubernetes GPU scheduling


Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Enrich IT Systems Administrator

Full Time United States $95000 - $110K per year Software Development

Senior Machine Learning Engineer, West Coast

Full Time Canada, United States 120K - 160K per year Software Development

Software Development Engineer, Backend (Go/Python)

Full Time Mexico Software Development

Fullstack QA Engineer (MSME)

Full Time United Kingdom Software Development

Identity Security Engineer

Full Time United States Software Development

Data Analyst

Full Time India Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 220,402+ Jobs in Software Development

Answer easy questions

Answer easy questions

220,402+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified