See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewQuestions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will design and implement AI evaluation frameworks, including defining ground truth, metrics, and scoring methods to measure model performance. Additionally, you will partner with engineering and product teams to operationalize these evaluations and drive system improvements based on your findings.
Who We Are:
Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.
Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.
Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.
Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.
Our Team Members:
We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!
We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.
Your Role: We're looking for a Senior Data Scientist, AI Evaluation to design how Alpaca measures whether our models and agents are actually right. You'll be a senior individual contributor who turns ambiguous quality questions into ground truth, scoring methods, and eval loops that the company can trust—and uses those results to make the systems better. You'll build on an established data foundation, so the focus is raising quality and speeding up safe rollout. You'll own the quality bar, independent of the teams that build and optimize those systems.
This role is for someone who cares as much about whether an answer is correct as about whether a model can generate one. You'll partner with Product, Engineering, Analytics Engineering, and business stakeholders to define what "good" looks like, build the evaluations that test it, and close the loop so evals drive iteration. If you have a strong quantitative background, have shipped rigorous, measurable work (evaluation, experimentation, or model validation), and want ownership over a greenfield eval practice at a fast-growing brokerage-infrastructure company, this is the role.
What You'll Do
What We're Looking For
Nice to Have
Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 221,257+ Jobs in Data Scientist
Answer easy questions
221,257+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”