For Employers

Anyone AI

Research Scientist (Remote/US/LATAM)

Posted 3 months ago
5-10 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

The role involves designing and owning frontier-grade evaluation packages for LLM capabilities across reasoning, coding, and multi-modal tasks. You will manage expert pools and act as the technical point of contact for labs to translate measurement needs into rigorous evaluation designs.

Research Scientist, Evaluations

Anyone AI Labs — Human Data Division
Reports to: CEO · Remote / LatAm / US

The role

You will own how Anyone AI measures frontier model capability. This is a research role at heart: you decide what a good evaluation is, design the benchmarks that prove it, and defend the methodology under lab scrutiny. You'll build frontier-grade evaluation packages across reasoning, coding, agents, tool use, and multi-modal — grounded in expert-verified truth, validated against multiple models, and QC'd to survive buyer-side review.

Responsibilities

  • Evaluation research. Turn public benchmarks and eval targets into original evaluation designs. Own the hard questions: construct validity, discrimination, headroom, and contamination.

  • Benchmark development. Build evaluation packages with subject-matter experts, each with expert-verified ground truth, multi-model headroom results, and rigorous QC (calibration layers, severity-weighted rubrics, deterministic verifiers).

  • Experts. Recruit, calibrate, and review a pool across coding, agentic/tool-use, and STEM/reasoning. Be the final arbiter of correctness and frontier difficulty.

  • Lab relationships. Be a technical point of contact for labs, with CEO support. Understand what they're trying to measure and translate it into an evaluation design.

  • Delivery. Turn lab requests into winning sample packages, then own pilots end to end. Nothing ships before it's lab-ready.

What we're looking for

  • Research background in ML evaluation or benchmarking — published/open benchmarks, eval research, or equivalent hands-on work labs relied on.

  • Deep LLM benchmarking expertise, with real strength in code-model evaluation.

  • Fluency with how frontier models are measured: rubrics, pass rates, headroom, contamination, and what makes a task discriminate a model.

  • Proven ability to hold a team or expert pool to a rigorous standard.

  • Fluent English. Spanish a nice to have.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Others jobs →

AI Response Evaluator

Freelance France, Japan, Mexico +3 more $10K – $20K Others

Admin Officer (Construction) - Work from Home / Midshift

Full Time Philippines Others

Intermediate Test Automation Engineer - OP02261

Full Time Brazil Others

L1 Technical Support Representative - Work from Home

Full Time Philippines Others

Vice President of Enterprise Account Management

Full Time United States Others

Clinical Provider (Chachoengsao, Thailand)

Part Time Thailand Others
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 219,827+ Jobs in Research Scientist

Answer easy questions

Answer easy questions

219,827+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified