For Employers

Intergral

AI Evaluation & Benchmarking Engineer

Posted 6 days ago
£60000 - £75000 per year
2-5 years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

You will design and build automated evaluation systems to measure the quality, reliability, and performance of AI-driven observability tools. This involves creating synthetic scenarios, benchmarking models, and collaborating with engineering teams to drive continuous product improvement.

AI Evaluation & Benchmarking Engineer

Location: Remote (United Kingdom)
Salary: £60,000–£75,000 DOE
Contract: Full-time, 40 hours per week
Employer: Intergral UK
Reports to: Director of Engineering

You must already have the right to work in the UK, as we're unable to sponsor visas for this role.


About the Role

We're looking for an AI Evaluation & Benchmarking Engineer to determine how we measure the quality of OpsPilot, and to build the systems that do it.

More important than any individual technology is how you approach measurement. We're looking for someone who questions whether a system is actually achieving the outcome it was designed for, works out how to measure that objectively, and builds what's needed to keep measuring it as the product changes.

You'll have real autonomy over the technical approach. We set the goals and check in regularly, but the expertise on how to get there is yours — we're hiring you because we need someone who can take an ambiguous problem and deliver a working system without being handed the steps. This is a hands-on engineering role with no reports and no QA silo. Our engineering structure is flat: you'll report to the Director of Engineering and work alongside the other engineers as a peer.

This isn't a traditional QA role. You won't be manually testing tickets, acting as a release gatekeeper or simply checking whether features technically work.


About OpsPilot

OpsPilot is an AI-led observability platform helping engineering and operations teams move from monitoring data to evidence-backed operational understanding — so they can investigate problems faster and act with greater confidence.

Intergral has more than 20 years of experience in application performance and observability, with an established customer base built around FusionReactor. OpsPilot is expanding beyond its historic Java and ColdFusion roots into the wider observability market, supporting OpenTelemetry-based metrics, logs and traces.

At the heart of OpsPilot is Coworker, an AI operations capability that continuously investigates telemetry, identifies situations that need attention and provides evidence-backed findings and recommended next steps.


What You'll Do

Evaluate the agent and the product it runs on

Coworker's findings are only as good as the system underneath them. An investigation can fail because the model reasoned badly, because retrieval surfaced the wrong evidence, because ingestion dropped a trace, or because the finding was presented in a way no engineer could act on. Evaluating the agent in isolation would tell us very little, so this role covers both.

On the agentic side, you'll:

  • Build automated evaluations for our agentic AI capabilities.
  • Create realistic synthetic scenarios, datasets and workloads with meaningful ground truth and evaluation criteria.
  • Measure task success, diagnostic accuracy, evidence quality, reliability, consistency, latency and cost.
  • Account for the non-deterministic nature of AI systems through repeated runs, variance analysis and determining whether changes are meaningful rather than noise.
  • Benchmark models, prompts, tools, retrieval strategies and agent workflows against repeatable baselines.
  • Use relevant industry benchmarks and standards, including established SRE practices and emerging AI-agent, AIOps and incident-response benchmarks, and build our own where existing approaches don't represent real operational work.

Across the wider product, you'll:

  • Extend evaluation across important customer journeys, APIs, backend services and UI.
  • Use OpenTelemetry, including its Semantic Conventions, to make benchmark environments representative and portable.
  • Use product telemetry and correlate benchmark results with metrics, logs and traces to understand why failures occur.
  • Build end-to-end measurements focused on customer outcomes rather than isolated components.

When we change a model, prompt, tool or agent workflow, we want to know what became better, what became worse and why — including the impact on quality, reliability, latency and cost.


Turn what we learn into continuous improvement

  • Turn failures and real-world problems into new evaluation scenarios.
  • Identify recurring failure patterns and capability gaps.
  • Test potential improvements against established baselines, holdouts and unseen scenarios.
  • Detect regressions, benchmark overfitting and improvements that don't generalise.

Longer term, this evaluation system becomes the harness for controlled self-improvement: identifying weaknesses, testing potential changes and objectively determining whether they should be retained. That's the direction of travel rather than the first year's work, but it's why we're building this properly.


Work with the rest of engineering

  • Work directly with AI, software, platform and SRE engineers to investigate findings and improve the product.
  • Make evaluation failures clear, reproducible and actionable.
  • Make straightforward fixes yourself where that's the most efficient approach.
  • Build tooling that makes evaluations easy for other engineers to create, run and understand.
  • Use AI-assisted engineering where it improves the speed or quality of your work.

Benchmarking should provide continuous feedback that helps engineering improve the product. This role is not a release gatekeeper.


What We're Looking For

Above everything else: the ability to take an ambiguous technical problem, develop an approach and deliver a working system independently.

We're more interested in demonstrated ability than an exact number of years, but we'd generally expect around 3+ years of relevant technical experience. Your background might be as an AI engineer, software engineer, SRE, platform engineer, performance engineer, SDET or similar.

Alongside that, we're looking for:

  • Practical experience working with LLMs, AI agents or AI evaluation.
  • Strong software engineering skills, particularly Python or a similar language.
  • Experience building automated evaluation, benchmarking, testing or experimentation infrastructure.
  • Experience creating synthetic workloads, datasets or evaluation scenarios.
  • An understanding of non-deterministic evaluation, including repeated measurement, variance and distinguishing meaningful changes from noise.
  • The ability to turn complex system behaviour into measurable criteria.
  • Comfort working across APIs, distributed systems and multiple layers of a software product.


Desirable

Any of the following would be useful, but none are required:

  • Agentic AI evaluation, tool use and multi-step workflows.
  • LLM evaluation frameworks and model-based evaluation techniques.
  • Automated experimentation or self-improving systems.
  • Dataset, ground-truth and holdout evaluation design.
  • Statistical experimentation and performance benchmarking.
  • OpenTelemetry, metrics, logs and distributed tracing.
  • SRE, incident response or observability.
  • Production SaaS and distributed systems.


What Success Looks Like

We have the beginnings of an evaluation harness, but the design and expertise are what we're hiring for.

We'd expect the first few months to go into the core evaluation harness and a starting corpus for Coworker's investigation quality, then extend outward across the rest of the product as that proves itself. How you sequence it is your call.

By six months, we should be able to objectively answer questions such as:

  • Is Coworker getting better at investigating operational problems?
  • Where does it perform well or poorly, and why?
  • Are its conclusions supported by the right evidence?
  • How do different models, tools and agent configurations compare?
  • What quality, latency and cost trade-offs are we making?
  • How do we compare against relevant external benchmarks?
  • Have improvements introduced regressions elsewhere?
  • Where should we focus improvement next?

Our evaluation corpus should keep growing as we encounter new problems.

Success isn't measured by the number of tests written or percentage test coverage. It's measured by our ability to understand how well OpsPilot is doing its job, where it isn't, and whether the changes we're making are actually making it better.


What We Offer

A small company rather than a large one, with the trade-offs that implies. Under ten people in engineering, a flat structure, and decisions made in a conversation rather than across three meetings. You'll have genuine influence over how this is done, and very little bureaucracy to work through to get there.

  • Fully remote within the UK.
  • Flexible working hours.
  • 25 days holiday plus bank holidays.
  • Real autonomy over your technical approach and how you deliver the role.


Our Interview Process

Straightforward: usually two or three conversations, with no technical coding tests.

If this sounds like the kind of challenge you're looking for, we'd love to hear from you.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Binance Accelerator Program - Backend Engineer (Web3)

Internship Italy, Taiwan Software Development

L2 Systems & Server Engineer

Full Time India 900K - 1200K per year Software Development

Senior SAP HR ABAP Developer, L3 Harris

Full Time Belgium $66100 - $102K per year Software Development

Operations Security Engineer – Remote-First

Full Time France, Germany, Spain Software Development

Principal Enterprise Applications SAP Public Cloud Consultant

Full Time Belgium $116K - $150K per year Software Development

CRL/SP AI and Data - Business Consulting

Full Time Belgium $210K - $256K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 214,734+ Jobs in Software Development

Answer easy questions

Answer easy questions

214,734+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified