Please mention DailyRemote when applying
Gray Swan is on a mission to empower the world to use AI safely and securely. We evaluate AI models for the leading frontier labs along with building real-time threat detection and adaptive adversarial red teaming agents for teams deploying AI.
We're a team of approximately 50 people, well-funded, growing quickly. Our work directly influences how the world deploys AI agents and systems at scale..
Come build and lead Gray Swan's cyber safety capability from the ground up, serving as the technical authority on AI-enabled cyber risk across red-teaming, evaluation, benchmarking, defenses, and safety infrastructure development. You'll help define how frontier AI systems are evaluated for offensive cyber capabilities while partnering with leading AI labs to reduce real-world security risks.
This role sits at the intersection of offensive security, AI safety, and machine learning. You'll transform deep cybersecurity expertise into scalable evaluation methodologies, safety infrastructure, and automated defenses that help establish industry standards for frontier model security.
If you have deep expertise in offensive cybersecurity, vulnerability research, or adversarial AI security, experience evaluating frontier models, and are driven to reduce catastrophic cyber risks from increasingly capable AI systems, we'd love to hear from you.
Design and lead adversarial evaluations of frontier LLMs for offensive cyber capabilities, including vulnerability discovery, exploit development, malware generation, privilege escalation, social engineering, persistence, and autonomous cyber operations across text, agentic, and multimodal systems.
Partner closely with machine learning engineers to translate cybersecurity expertise into scalable benchmarks, classifiers, guardrails, automated detection systems, and evaluation infrastructure for both internal products and frontier AI lab deployments.
Develop and maintain Gray Swan’s catastrophic cyber harm taxonomy, continuously evolving cyber evaluation frameworks as frontier model capabilities rapidly advance.
Produce technical risk assessments and actionable recommendations for frontier AI labs, enterprise customers, and internal stakeholders, helping guide responsible model deployment and security mitigations.
Build, mentor, and lead a world-class team of cybersecurity subject matter experts while establishing scalable evaluation processes, quality standards, and technical infrastructure.
Represent Gray Swan as the company's cybersecurity authority, collaborating with frontier AI labs, security researchers, government partners, and the broader AI safety and cybersecurity communities.
Deep technical expertise in offensive cybersecurity, vulnerability research, exploit development, penetration testing, malware analysis, reverse engineering, or a closely related field through industry, research, or equivalent experience.
Significant experience assessing advanced cyber threats, offensive tooling, or AI-enabled cyber capabilities, especially in critical infrastructure domains.
Hands-on experience conducting adversarial evaluations, AI red-teaming, LLM security research, or building evaluation datasets for frontier AI systems.
Comfortable operating at the intersection of cybersecurity research, AI safety, and machine learning engineering.
Thrive in highly ambiguous, fast-moving environments where you'll define strategy while building entirely new capabilities.
A builder who enjoys creating teams, infrastructure, and evaluation systems from scratch.
Experience developing machine learning models, AI security classifiers, or automated cyber detection systems.
Hands-on experience red-teaming frontier language models, jailbreaking, prompt injection research, or agentic AI evaluations.
Experience working with frontier AI labs, national security organizations, or leading cybersecurity research teams.
Background in threat intelligence, autonomous cyber operations, AI agent security, or AI governance.
Strong software engineering experience in Python, Go, Rust, or other systems programming languages.
If you don’t have 100% of these, you should still seriously consider applying. We care more about what you can do than your credentials.
Want visibility across the frontier of AI security by evaluating multiple frontier models and working directly with the world's leading AI labs.
Are excited to build the infrastructure that makes AI cyber safety scalable.
Feel motivated to reduce catastrophic cyber risks posed by increasingly capable AI systems.
Enjoy turning offensive security research into practical defenses that improve the safety of frontier AI.
Have a vision for what world-class AI cyber safety should look like—and are excited to build it.
We offer a competitive compensation package designed to reward impact and incentivize growth. Our compensation philosophy is informed by our current valuation and recent industry data.
Compensation: $230,000 - $280,000 plus performance based bonus and meaningful equity package
Benefits:
401k with up to 4% matching
28 days annual leave (vacation + holidays)
Health, dental, and vision coverage
Catered lunches (Pittsburgh office)
Flexible work arrangements
Visa sponsorship available for exceptional candidates
🔎 Application review. We read everything; we’ll respond within 10 days.
✏️ Online technical screen (15 min). Complete a simple, job-relevant exercise.
🗣 Intro call (30 min). We learn about you; you learn about us.
🧑💻 Technical interview (90 min). Live coding with some tasks requiring AIand others not.
🗣 Experience & culture interview (60 min). Conversational exploration of the skills fit.
😇 Reference checks. We’ll reach out to 3-5 references that you provide.
📃 Offer. If it’s mutual, we move fast.
Submit your resume, link to your portfolio, and answer the questions on the application.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Others
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!