United States$200K - $350K per year2-5 yrs expOthers
You will design and implement post-training systems, including fine-tuning, reinforcement learning, and evaluation frameworks for frontier AI models. You will also collaborate with researchers and domain experts to build scalable data pipelines and turn experimental findings into durable infrastructure.
The AI Red Teamer will stress-test large language models by designing adversarial prompts to identify vulnerabilities such as bias, hallucinations, and safety guardrail failures. They will also document experiments, score model responses against harm taxonomies, and collaborate with researchers to strengthen AI defenses.
You will design adversarial prompts and multi-turn interaction chains to test AI models for vulnerabilities related to cyberattacks and malware generation. Additionally, you will evaluate model-generated technical output for functional correctness and document findings within a structured cybersecurity framework.
Design and execute adversarial prompts to test AI model safety against CBRNE threats while documenting failure points. Collaborate with researchers and policy teams to refine evaluation frameworks and improve model defenses.