Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
About Gramian
Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.
Role Overview
We are looking for a materials science professional to contribute to an advanced AI research project focused on building rigorous STEM coding datasets. The role combines materials science expertise, Python-based scientific computing, and AI evaluation, with a focus on creating well-structured scientific problems, implementing verified solutions, and developing tests that accurately distinguish correct from incorrect model outputs.
Responsibilities
Design scientific coding tasks with one main problem and at least 3 logically connected sub-problems.
Implement verified Python solutions with complete unit test coverage.
Create discriminative test cases that distinguish correct from incorrect AI-generated outputs.
Perform quality control checks using the Central Task Platform (CTP), including Tier 1 structure checks and Tier 2 quality rubrics.
Revise tasks and solutions based on QC feedback.
Optimize tasks against Pass@K evaluation criteria across multiple LLM judges, including GPT, Gemini, and Nemotron.
Validate scientific correctness, determinism, well-posedness, and expected outputs.
Maintain a low rework rate and high first-submission quality.
Participate in project reviews, feedback sessions, and standups during required overlap hours.
CONTRACT: Freelance / Contractor
COMMITMENT: Full-time commitment; overlap requirements to be confirmed
LOCATIONS: Remote; eligible locations to be confirmed
PROCESS: Not specified
Requirements
Design scientific coding tasks with one main problem and at least 3 logically connected sub-problems.
Implement verified Python solutions with complete unit test coverage.
Create discriminative test cases that distinguish correct from incorrect AI-generated outputs.
Perform quality control checks using the Central Task Platform (CTP), including Tier 1 structure checks and Tier 2 quality rubrics.
Revise tasks and solutions based on QC feedback.
Optimize tasks against Pass@K evaluation criteria across multiple LLM judges, including GPT, Gemini, and Nemotron.
Validate scientific correctness, determinism, well-posedness, and expected outputs.
Maintain a low rework rate and high first-submission quality.
Participate in project reviews, feedback sessions, and standups during required overlap hours.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 212,664+ Jobs in Others
Answer easy questions
212,664+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”