For Employers

Vetto AI

AI Research

Posted 4 days ago
Worldwide
2-5 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

You will design and execute experiments to evaluate model behavior, create benchmarks, and develop training strategies for AI models. You will collaborate with AI labs to turn research goals into concrete data projects and technical reports.

AI Researcher

About Vetto

Vetto builds the infrastructure for next-generation AI training data. We partner with the world’s top AI labs to push the frontier of what models can do — by designing, collecting, and delivering the highest-quality training data for the hardest problems in AI.

AI Researchers at Vetto figure out how to teach models new capabilities and how to tell whether it worked. You’ll design datasets and evaluations, run experiments, study model failures, and turn what you learn into better data and training strategies.

Most of our work is evaluation-driven, with some fine-tuning and post-training used to validate research hypotheses. You won’t be expected to build the infrastructure alone: you’ll work with an engineering team that builds the platforms and tools behind the research. You should still be comfortable writing code, analyzing data, and moving your own experiments forward.

Why Vetto?

  • End-to-end ownership. You’ll take research questions from an initial idea to a dataset, benchmark, experiment, and useful conclusion.
  • Direct collaboration with top AI labs. Your work will help frontier teams understand their models and decide what data or training approach to try next.
  • Work that gets used. Your research can become training data, evaluations, technical reports, open benchmarks, and published papers.
  • A wide range of hard problems. Depending on the project, you might work on agents, model evaluations, human or synthetic data, failure analysis, or post-training.
  • Fast, flat team. You’ll ship experiments in days, not quarters.

What You’ll Do

  • Design and run experiments to understand model behavior, test new ideas, and determine whether a dataset or training intervention actually works.
  • Use fine-tuning and post-training experiments when useful to validate research hypotheses and measure improvements in the capabilities we care about.
  • Create datasets, benchmarks, agent environments, rubrics, and evaluation methods for problems that are not well measured today.
  • Work directly with AI labs and enterprise partners to turn open-ended model-development goals into concrete research and data projects.
  • Dig into model outputs and trajectories to understand systematic failure modes and separate model limitations from problems in the data, grader, or evaluation setup.
  • Write analysis code and build lightweight tools or prototypes for your own work, with support from engineers when a workflow needs to become reliable infrastructure.
  • Share what you learn through clear technical reports, internal discussions, benchmarks, and research papers.

What We’re Looking For

Must-haves:

  • Experience doing empirical research or similarly open-ended technical work with machine learning systems.
  • Hands-on experience designing and running agentic systems for complex work, such as coding agents and automated research pipelines.
  • Good experimental judgment: you can form a hypothesis, choose useful baselines and metrics, run a careful experiment, and make sense of noisy results.
  • Strong Python and data analysis skills, plus enough software engineering ability to prototype your own research workflows.
  • The independence to take an ambiguous question about a model and turn it into a dataset, benchmark, experiment, or other concrete contribution.
  • Rigorous analytical thinking — you look at the underlying evidence, notice subtle failure modes, and question results that seem too simple or too good to be true.
  • Excellent written and verbal communication with researchers, engineers, customers, and non-specialists.
  • Comfort with ambiguity, fast iteration, changing priorities, and occasional forward-deployed work with partner teams.

Nice to have:

  • Research or industry experience in LLM evaluation, post-training, agents, human or synthetic data, alignment, or AI safety.
  • Experience designing datasets or benchmarks, validating annotation quality, or working with large collections of model outputs and trajectories.
  • Published research, technical writing, open-source contributions, or strong independent projects.
  • Experience building or managing human feedback pipelines.
  • A background in cognitive science, linguistics, philosophy, statistics, or another field that helps you think clearly about evaluation.
  • Experience at an AI lab, AI data company, or in a partner-facing research role.

We care more about evidence of excellent work than any single credential. A graduate degree, publications, lab experience, open-source work, independent research, and production engineering experience are all useful signals, but none is a requirement on its own.

What We Offer

  • Competitive compensation and a stock option plan.
  • Global, remote-first work with flexibility and periodic in-person on-sites every few months.
  • Direct collaboration with cutting-edge AI labs on evaluation, data, and post-training challenges.
  • High ownership and fast career growth in a founding-era team.

Location: Global / Remote, with periodic in-person on-sites every few months

Reports to: Co-founders

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Backend Engineer - Agentic AI (m/w/d)

Full Time Germany, United States €60000 - €80000 per year Software Development

Solutions Engineer - GVSE Mid-Market

Full Time United States $155K - $203K per year Software Development

AI Experience Designer

Full Time United States $120K - $135K per year Software Development

Salesforce Architect All Levels

Full Time United States $75 - $145 per hour Software Development

Salesforce Quality Assurance All Levels

Full Time United States $38 - $67 per hour Software Development

Salesforce Business and Functional Consultant All Levels

Full Time United States $42 - $81 per hour Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 218,987+ Jobs in Software Development

Answer easy questions

Answer easy questions

218,987+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified