You will establish a distinct visual identity and brand presence for Inferact and vLLM while designing intuitive developer product experiences. You will collaborate with founders and engineers to build scalable design systems across web surfaces, internal tooling, and enterprise platforms.
You will work at the core of vLLM to optimize how models execute across diverse hardware and architectures. Your contributions will directly impact the performance and scalability of AI inference engines globally.
You will design and maintain the operational backbone for vLLM, including cluster management, deployment automation, and production monitoring. Your work will ensure that AI model inference systems are observable, debuggable, and highly reliable at a massive scale.
United States$200K - $400K per year5-10 yrs expOthers
You will own and operate high-performance GPU compute infrastructure to ensure it remains healthy, available, and scalable for engineering teams. This involves managing cluster health, GPU availability, monitoring, alerting, and incident response across various compute providers.
Lead the end-to-end execution of conferences, partner events, and GTM motions to build brand awareness for vLLM and Inferact. Develop strategic relationships with corporate partners and coordinate technical marketing assets and landing pages.
United States$200K - $400K per year5-10 yrs expOthers
Build and optimize AMD GPU backends, kernels, and runtime paths to make vLLM a first-class inference engine. Improve performance-critical paths including attention, GEMM, and communication-heavy operations using ROCm and related tooling.
United States$200K - $400K per year5-10 yrs expOthers
Optimize the vLLM inference engine to improve the speed and cost of running LLMs and diffusion models. Develop innovations for diverse hardware and architectures, including mixture-of-experts and multimodal models.
United States$200K - $400K per year5-10 yrs expOthers
Build and optimize AMD GPU and TPU backends, kernels, and compiler integrations to make vLLM a first-class inference engine on non-NVIDIA hardware. Improve critical paths such as attention, GEMM, and KV-cache while developing robust benchmarking infrastructure.
The role involves creating high-quality technical content, tutorials, and demos to help developers adopt and scale vLLM. You will act as an educator-builder, explaining complex inference systems concepts and hosting workshops for the AI infrastructure community.