You will establish a distinct visual identity and brand presence for Inferact and vLLM while designing intuitive developer product experiences. You will collaborate with founders and engineers to build scalable design systems across web surfaces, internal tooling, and enterprise platforms.
Inferact
7 Remote Job Openings at Inferact
Lead the end-to-end execution of conferences, partner events, and GTM motions to build brand awareness for vLLM and Inferact. Develop strategic relationships with corporate partners and coordinate technical marketing assets and landing pages.
Maintain and scale the compute infrastructure powering CI, releases, and performance benchmarks for the vLLM project across various accelerators. Focus on reducing CI time-to-signal and building tooling to support thousands of open-source contributors.
Member of Technical Staff, AMD GPU Performance Engineering
Inferact
·
Full Time
·
2 months ago
Inferact
Build and optimize AMD GPU backends, kernels, and runtime paths to make vLLM a first-class inference engine. Improve performance-critical paths including attention, GEMM, and communication-heavy operations using ROCm and related tooling.
Optimize the vLLM inference engine to improve the speed and cost of running LLMs and diffusion models. Develop innovations for diverse hardware and architectures, including mixture-of-experts and multimodal models.
Member of Technical Staff, TPU & AMD GPU Performance Engineering
Inferact
·
Full Time
·
2 months ago
Inferact
Build and optimize AMD GPU and TPU backends, kernels, and compiler integrations to make vLLM a first-class inference engine on non-NVIDIA hardware. Improve critical paths such as attention, GEMM, and KV-cache while developing robust benchmarking infrastructure.
The role involves creating high-quality technical content, tutorials, and demos to help developers adopt and scale vLLM. You will act as an educator-builder, explaining complex inference systems concepts and hosting workshops for the AI infrastructure community.