For Employers

Improbable

AI Researcher - Bolter

Posted 22 days ago
2-5 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

You will design and execute experiments to evaluate the reliability and capability of AI agents, turning research findings into actionable product decisions. Additionally, you will build evaluation layers and prototype research concepts into functional product features.

About Bolter

Bolter is an AI agent platform built to help people get real work done, without needing to stitch together multiple tools or spend weeks setting things up.

You describe what you need in plain language and Bolter creates and runs agents that can carry out the work end to end. They retain context, remember how you work and can also create real, shareable apps such as trackers and dashboards.

Bolter is funded and currently at the validation-sprint stage, working with a small group of high-impact operators. The product already exists. We are now looking for a founding designer to lead a significant redesign and establish the design foundations for what comes next.

 

About the role

As Bolter's AI Researcher, your job is to make agents that do real work trustworthy, capable, and useful - and to figure out how before we build it at scale.

This is applied research with direct product impact. You design and run experiments, evaluate what works and what doesn't, and turn findings into decisions the engineering team ships. You'll work directly with the GM and the product engineer in a lean team, with specialised AI agents supporting implementation.

The research problem goes beyond improving a model. It is working out how an agent that retains context, carries out multi-step work and creates software itself can be relied on to do it correctly.

 

What you’ll do

  • Design and run experiments to test how Bolter's agents behave - reliability, context retention, multi-step task completion, and where they fail.

  • Build and own the evaluation layer. Design evals that measure whether agents actually do the work, not just whether they sound plausible.

  • Research the frontier. Keep Bolter current on the state of the art in LLM agents, tool use, and reliability - and translate that into what we should build.

  • Turn findings into decisions. You produce clear, actionable recommendations the engineering team can ship, not just papers.

  • Prototype research into product. Take promising ideas from experiment to working prototype, and hand off what proves useful.

  • Shape the research roadmap as Bolter grows.

Why you're made for this

  • A track record of applied AI/ML research - ideally in LLM application design, agentic systems, evals, or production AI reliability.

  • Strong experimental design and analysis. You know how to test a hypothesis rigorously and read the result honestly.

  • Strong engineering fundamentals - you can prototype your own experiments, not just direct others to run them.

  • Deep familiarity with LLMs, tool use, context management, and the failure modes of agentic systems.

  • Strong written communication. Research that isn't understood and acted on doesn't help.

  • You're rigorous but pragmatic. You know when a finding is strong enough to act on and when it needs more evidence.

  • You're comfortable with ambiguity. Many of the problems we're solving don't have textbook answers.

  • You can leverage AI agents as a force multiplier - we run lean, with a small human core augmented by a fleet of specialised agents.

  • You care about craft, but you ship.

Bonus points

  • Experience with LLM evals, hallucination mitigation, or production AI reliability at scale.

  • Experience building or evaluating agentic systems, tool use, or autonomous workflows.

  • A public track record - papers, open source, blog posts, side projects. We love researchers who ship outside work too.

  • Experience in product-led research - where the output is a shipped feature, not just a finding.

While we think the above experience could be important, we're keen to hear from people who believe they have valuable experience to bring to the role. If you identify with the team and mission, but not all of our requirements, then please still apply!

Improbable Candidate Privacy Policy

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

EDA Customer Success Application Engineer

Full Time United States Software Development

Principal Data Engineer with GCP

Freelance Poland 22500 - 34500 per month Software Development

Software Engineer III

Freelance United States Software Development

Client Onboarding Lead, Cloud Accounting

Full Time Canada 75000 - 85000 per year Software Development

Service - Field Service Engineer (FSE) – Diagnostic Imaging (Vascular, Surgery X-Ray & Mammography) - UK Midlands (Birmingham)

Full Time United Kingdom Software Development

Certified Payroll Administrator

Freelance United States $40 - $45 per hour Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 219,685+ Jobs in Software Development

Answer easy questions

Answer easy questions

219,685+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified