Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
You will design complex AI prompts to test model limitations and create grading rubrics to evaluate AI performance. The process involves screen-recording your thought process while you refine prompts based on system failure points.
We're running a paid study to build a bench of people who are exceptionally good at designing tasks that expose AI model limitations. By turning real-world workflows into demanding requests, we can better evaluate where current models break down. This initial trial helps us identify individuals suited for ongoing prompt engineering and evaluation work.
You will spend about an hour translating a complex workflow from your job or personal life into a demanding prompt that requires reasoning and real-world lookup. After running it in ChatGPT to identify where the model fails, you will refine the prompt until it breaks the system. Finally, you will write a clear grading rubric that a stranger could use to evaluate any AI's attempt at your task. This entire process is screen-recorded, as we are assessing your thought process just as much as the final submitted files.
We welcome professionals, domain experts, and power users who have deep knowledge of specific workflows. You need to be capable of evaluating an AI's output within seconds and comfortable working on a laptop or desktop with a ChatGPT account. Candidates who excel at this trial will be considered for a long-term bench of evaluators.
Pick a familiar workflow and convert it into a demanding AI prompt
Test your prompt in ChatGPT to find failure points, making it harder if the AI succeeds
Write a comprehensive rubric for grading the AI's performance
Share your screen, camera, and microphone while completing the task
Submit your prompt, failure notes, rubric, and the generated output file
Deep familiarity with a specific professional or personal workflow
Ability to quickly evaluate the accuracy and quality of AI outputs
Access to a laptop or desktop computer
An active ChatGPT account
Comfortable being screen-recorded while thinking through complex tasks
$20 one-time
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Learn more at terac.com or on YouTube at @jointerac.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Development
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”