Please mention DailyRemote when applying
See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewUpload your resume and we draft a letter for this exact role, tailored to what it asks for.
Create original research tasks that evaluate AI systems by synthesizing clues from reputable open web sources. Document definitive answers, verify all sources, and refine tasks based on model performance and reviewer feedback.
Careerflow Human Data Labs partners with AI companies to bring real-world professional expertise into their products.
About the role
You'll create research tasks that evaluate how well AI systems can find verifiable answers across the open web. Each task requires synthesizing multiple independent clues from public English sources, each supported by exact quotes. Your work is tested against frontier models and reviewed for rigor—only tasks with bulletproof sources and appropriate difficulty count. Remote, contract position; approximately four hours per completed task.
What you'll do
Craft a 60–120 word prompt that feels like a genuine question—original phrasing, no copied content, no search hints.
Research four to six clues across different reputable open English sources (excluding Wikipedia and paywalled content), each with a direct quote, URL, and verification date.
Identify a definitive answer (name, date, catalogue number, or documented title) and clearly explain why alternatives don't fit.
Document plausible wrong answers and which clue eliminates each; articulate your search path.
Verify all sources while logged out; refine tasks that models solve too easily or that reviewers find unclear.
Required skills
3+ yrs of experience with Masters Degree
Advanced web research: search operators, reverse lookups, direct access to institutional catalogues and databases.
Familiarity with archives and catalogues: libraries, museums, legislative and court records, registers, specialized databases (sports, film), Wayback Machine, preprint servers.
Strong source evaluation: distinguishing primary from secondary sources, open from restricted, original from republished, stated facts from implications.
Precise citation: pulling verbatim text from web pages, PDFs, scans, and tables—typos included.
Clue architecture: designing constraints that are individually searchable yet jointly pinpoint the answer; identifying logical chains, information leaks, and undefined terms.
Clear, fluent written English.
Comfort using search engines and LLM tools—with a commitment to fact-checking every AI-generated claim against the actual source.
Must-have
Demonstrated sourced research background: journalism or fact-checking, library/archival work, genealogy, due diligence, graduate-level primary-source research, or citation-heavy editing.
Availability: 20+ hours per week, any time zone; NDA required (materials are confidential client IP).
Nice-to-have
Deep expertise in history, science/technology, sports, music, politics, geography, film, or television.
OSINT or cataloguing experience.
Prior work creating benchmarks, exams, or puzzle content.
3+ years of professional experience
Total number of positions: 320
Engagement Length: 4 weeks, Full-time (8 hours/day) 40 hours per week. Overlap 4 hours with PST
Location - Bangladesh, Brazil, Colombia, Egypt, Nigeria, Ghana, India, Indonesia, Turkey, Vietnam
6-week project, with a preferred commitment of up to 40 hours per week.
Fully remote, with payment in USD.
Start as soon as you successfully pass the assessment.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 215,081+ Jobs in Writing
Answer easy questions
215,081+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”