Please mention DailyRemote when applying
Staff Technical Program Manager, Deployments
Location: Remote
Employment Type: Full time
Location Type: Remote
Department: Infrastructure & Deployments
Evergrid is building the inference neocloud — frontier-level AI infrastructure powered by emerging hardware platforms for what we believe is the fourth industrial revolution. We design, deploy, and operate large-scale AI systems for research labs, hyperscalers, enterprises, and institutions, supporting training, inference, and HPC workloads at gigascale.
Our platform spans Private Clusters for customers with strict privacy and compliance requirements, a Compute Platform of managed Kubernetes, Slurm, networking, and storage, and Fleet & Infrastructure operations that keep gigascale AI factories provisioned, monitored, and healthy. With over 1GW of capacity in active development, a 24/7 engineering team held to a 12-minute response SLA, and partnerships with leading hardware manufacturers, we deliver stable capacity on schedule and at predictable cost.
We're looking for problem-solving, opportunity-finding teammates who operate with high conviction and high trust — people who see this as a chance to set the standard of excellence for the industry, not just fill a seat in it. If you want to do the most meaningful work of your career and help build the infrastructure the next generation of AI runs on, come build with us at Evergrid.
Evergrid's TPM function is still being built, which means there is a real opportunity to shape how deployments run rather than inherit how they already work.
We are hiring a Staff TPM who will own deployment programs — standing up new capacity, expanding existing sites, and delivering infrastructure for customer clusters. Our TPMs define and drive the entire deployment engagement model, from hardware vendor engagement through first customer cluster delivery.
If you have spent your career running infrastructure deployment programs at a hyperscaler or a leading neocloud, understand GPU and accelerator architecture, and have shaped deployment frameworks rather than just operated within them, this is the role for you.
• Own the infrastructure deployment for new capacity or site expansions end-to-end: hardware vendor and OEM dependencies, architecture updates, compute platform foundations work, commissioning gate framework definition, and first customer cluster delivery as the success metric.
• Lead Deployment Phase 0 on the TPM side: define firmware version targets before kickoff, set commissioning gate criteria, and define the DRI matrix.
• Manage compounding cross-hardware-generation dependencies where active production programs run in parallel with capacity expansion projects, and prevent them from competing for the same engineering pool without a plan.
• Own real-time execution dashboards; deliver crisp, data-driven executive updates that surface decision elements without requiring follow-up.
• Govern cross-organizational dependencies without waiting for escalation authority.
• Coach more junior TPMs on technical depth, risk identification, and executive communication.
• Actively drive AI tool integration across your programs; identify where AI materially improves program tracking, risk detection, and executive communication.
• Deep, working fluency with GPU and accelerator architecture across hardware generations, firmware lifecycle (driver stacks, BIOS/BMC), compute orchestration, SDN, storage, networking (leaf-spine topology, ZTP, fabric commissioning), and monitoring/observability.
• Direct hardware partner engagement: personal ownership of GPU or OEM certification and validation timelines, not coordination feeding into someone else's relationship.
• Active daily use of AI tools to drive program-level outcomes: risk detection, dependency mapping, data analysis, and executive communication, not just personal productivity.
• 7+ years as a Technical Program Manager with a track record of owning infrastructure deployment programs end-to-end at a hyperscaler, GPU cloud provider, or AI infrastructure company, ideally with direct experience in large-scale capacity expansion projects.
• Proven ability to define deployment engagement models from scratch, not just operate within existing frameworks, and make them stick across engineering organizations that didn't ask for them.
• Track record of driving cross-organizational alignment at VP/SVP level without formal authority, including building durable alignment on programs that fall in the cracks between teams.
• Exceptional written and verbal communication for delivering clear, data-driven, decision-oriented updates to executive stakeholders.
• Experience defining or substantially redesigning a commissioning gate framework for site deployments.
• Experience coaching or developing more junior TPMs in technical depth and program execution.
• Competitive compensation and equity packages
• Paid time off, paid holidays & leave of absence programs
• Comprehensive health, dental & vision insurance
• Home office stipend
• 401(k) Retirement plan with company match up to 4% of salary
Compensation will be paid in the range of $200,000 to $240,000 + Bonus. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.
Evergrid is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Technical Program Manager
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!