Member of Technical Staff, Inference
You will work at the core of vLLM to optimize how models execute across diverse hardware and architectures. Your contributions will directly impact the performance and scalability of AI inference engines globally.