For Employers

Inferact

Remote Job Openings at Inferact (10)

Member of Technical Staff, Inference

United States 2-5 yrs exp Others

You will work at the core of vLLM to optimize how models execute across diverse hardware and architectures. Your contributions will directly impact the performance and scalability of AI inference engines globally.

Product Marketing Manager

United States 2-5 yrs exp Product

Lead the end-to-end execution of conferences, partner events, and GTM motions to build brand awareness for vLLM and Inferact. Develop strategic relationships with corporate partners and coordinate technical marketing assets and landing pages.

Member of Technical Staff, AMD GPU Performance Engineering

United States $200K - $400K per year 5-10 yrs exp Others

Build and optimize AMD GPU backends, kernels, and runtime paths to make vLLM a first-class inference engine. Improve performance-critical paths including attention, GEMM, and communication-heavy operations using ROCm and related tooling.

Member of Technical Staff, Inference

United States $200K - $400K per year 5-10 yrs exp Others

Optimize the vLLM inference engine to improve the speed and cost of running LLMs and diffusion models. Develop innovations for diverse hardware and architectures, including mixture-of-experts and multimodal models.

Member of Technical Staff, TPU & AMD GPU Performance Engineering

United States $200K - $400K per year 5-10 yrs exp Others

Build and optimize AMD GPU and TPU backends, kernels, and compiler integrations to make vLLM a first-class inference engine on non-NVIDIA hardware. Improve critical paths such as attention, GEMM, and KV-cache while developing robust benchmarking infrastructure.

Member of Technical Staff, Developer Relations

United States $200K - $400K per year 5-10 yrs exp Software Development

The role involves creating high-quality technical content, tutorials, and demos to help developers adopt and scale vLLM. You will act as an educator-builder, explaining complex inference systems concepts and hosting workshops for the AI infrastructure community.