See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewUpload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will own the evaluation framework for AI systems by constructing ground truth datasets, designing audit methodologies, and measuring model performance. You will provide independent, objective evidence to support release decisions while ensuring system trustworthiness and fairness.
ABOUT DEFCON AI
RESILIENCE IN THE FACE OF DISRUPTION. DEFCON AI is an insights company that leverages artificial intelligence, mathematical optimization, data analytics, and software engineering for resilient optimization of complex systems.
In today’s dynamically changing world, DEFCON AI’s technology aligns outcomes with operational goals, better decision making, and empowers customers to anticipate assess, and mitigate the impacts of disruptions.
Be the independent voice that keeps the whole program honest — your evaluation is the standard everyone else is held to.
About the Role
You'll join the analytics and AI engineering team behind a system that genuinely matters: an AI-assisted platform that brings together records from dozens of disparate data sources, resolves them to the correct individual, highlights what analysts should review first, and provides transparent, explainable recommendations that users can trust. Operating within a secure government cloud environment, the platform tackles complex challenges in AI, data integration, and decision support where quality, trust, and accountability are mission-critical.
As a Model Test & Measurement Engineer, you'll own the evaluation framework that helps ensure those systems perform as intended. You'll build and maintain labeled ground truth datasets, design statistically sound audit and sampling methodologies, measure model performance across releases, and create the evidence packages that support deployment decisions. You'll independently validate both scoring and generative AI capabilities, helping the team understand not only whether a model works, but how confidently its outputs can be trusted.
This is a role with genuine influence. Your assessments will inform release decisions, drive improvement efforts, and provide the objective evidence customers rely on when evaluating system performance. You'll work closely with data scientists, AI engineers, and technical leadership while maintaining the independence needed to provide clear, defensible evaluations. You do not build the models you validate, and you do not approve thresholds or release authorization against your own evidence.
If you're energized by measurement, validation, and making complex AI systems more trustworthy, this is an opportunity to have an outsized impact on both the technology and the mission it supports.
This is a fully remote role with occasional travel to DEFCON AI headquarters, customer sites, and partner facilities as needed.
Key Responsibilities
Required Qualifications
Preferred Qualifications
What Success Looks Like
What We Offer
Salary Range: $150,000–$190,000. This represents the typical salary range for this position based on experience, skills, and other factors.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 212,134+ Jobs in Software Development
Answer easy questions
212,134+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”