For Employers

FindErnest

Quality Engineering Consultant - AI/ML Testing

Posted 3 days ago
Worldwide
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

The consultant will design and execute comprehensive test strategies for AI and GenAI applications, ensuring accuracy, safety, and performance. They will collaborate with cross-functional teams to implement automation frameworks and establish quality metrics for LLM-based solutions.

This is a remote position.

We are seeking a highly motivated Quality Engineering (QE) Consultant with 4 - 6 years of experience in software testing, automation, and emerging AI technologies. The ideal candidate will have hands-on experience testing AI/GenAI applications, APIs, and web applications, along with strong expertise in Playwright automation and AI evaluation frameworks.

This role will focus on ensuring the quality, reliability, safety, and performance of AI-powered solutions by designing comprehensive test strategies, implementing automation frameworks, validating AI model behavior, and establishing evaluation mechanisms for large language model (LLM)-based applications.

The candidate will collaborate closely with Product Owners, Developers, Data Scientists, and AI Engineers to drive quality throughout the software development lifecycle.

AI & GenAI Testing :

- Design and execute test strategies for AI/GenAI applications, including LLM-powered solutions, AI agents, copilots, and conversational systems.

Validate AI outputs for :

1. Accuracy

2. Relevance

3. Groundedness

4. Consistency

5. Hallucination detection

6. Toxicity and safety compliance

- Create and maintain prompt test suites and benchmark datasets.

- Perform functional, regression, performance, and reliability testing for AI-enabled applications.

- Define and execute AI evaluation metrics and quality gates.

Automation & Quality Engineering :

- Develop and maintain test automation frameworks using Playwright.

- Automate UI, API, and end-to-end business workflows.

- Create reusable automation assets and testing accelerators.

- Integrate automated tests into CI/CD pipelines.

- Support shift-left testing and quality engineering practices.

API Testing :

- Design and execute API test cases using tools such as :

1. Postman

2. REST Assured

3. Playwright API Testing

- Validate API contracts, authentication, authorization, error handling, and performance.

- Conduct integration testing across distributed systems and third-party services.

AI Evaluation & Quality Metrics :

- Establish AI evaluation frameworks and quality scorecards.

- Measure model quality using evaluation techniques such as :

1. Precision/Recall

2. Relevancy scoring

3. Semantic similarity

4. Groundedness validation

5. Human-in-the-loop evaluation

- Analyze AI quality trends and recommend improvements.

- Support Responsible AI and model governance requirements.

Collaboration & Stakeholder Management :

- Work closely with engineering, product, and AI teams to identify quality risks early.

- Participate in requirement reviews and design discussions.

- Communicate testing progress, risks, and quality metrics to stakeholders.

- Contribute to QE best practices, standards, and reusable playbooks.

Requirements

Experience :

- 4 - 6 years of experience in Quality Engineering, Software Testing, or Test Automation.

- Hands-on experience testing AI/Generative AI applications.

- Experience with modern web application testing and API testing.

Technical Skills :

- Strong expertise in Playwright automation.

- Hands-on experience with API testing and automation.

- Experience validating AI/LLM-based applications and AI agents.

- Good understanding of AI evaluation methodologies and testing techniques.

- Familiarity with :

1. M365 Copilot

2. OpenAI/Azure OpenAI

3. Claude

4. Gemini

5. Agentic AI frameworks

- Knowledge of test management tools such as Azure DevOps (ADO), Jira, or similar platforms.

- Experience working with CI/CD pipelines and DevOps practices.

Preferred Skills :

- Exposure to Python or TypeScript for automation and AI testing.

- Experience with AI evaluation frameworks such as :

1. RAG evaluation

2. Prompt evaluation

3. LLM bench-marking

4. Human feedback-based evaluation

- Knowledge of Responsible AI principles, AI governance, and model risk management.

- Experience testing Retrieval-Augmented Generation (RAG) systems and AI agents.

Education / Professional qualifications :

- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Vendor Growth Analytics Sr Analyst

Full Time Argentina Software Development

Senior UC Engineer

Full Time United States Software Development

Quantum Computing Patent Agent

Full Time United States $120K - $180K per year Software Development

Staff Full-Stack Engineer, Experience Team

Full Time United States $149K - $262K per year Software Development

Android Developer

Full Time Israel Software Development

Signal & Power Integrity Hardware ASIC Simulation Engineer Lead (Remote)

Full Time United States $183K - $263K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Software Development

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified