For Employers

Anyone AI

Senior Software Engineer – Open Source & SWE-Bench Evaluation

Posted 21 days ago
$65 per hour
2-5 years experience
Apply Now

Please mention DailyRemote when applying

?
Resume Match Score

See how much of this job your resume covers, and what’s missing.

Want a recruiter to go through it line by line?

Get professional review

Create a cover letter for this job

Upload your resume and we draft a letter for this exact role, tailored to what it asks for.

  • Tailored to this role
  • Based on your resume
  • Fully editable
AI Summary

You will review and evaluate machine learning challenges to ensure they are technically sound, reproducible, and appropriately difficult. This involves analyzing datasets, metrics, and pipelines to detect flaws like data leakage, label noise, and metric gaming.

Anyone AI is recruiting experienced Machine Learning Engineers for a specialized project focused on reviewing and evaluating machine learning challenges used in AI model training and evaluation.

The work involves analyzing ML experiments, datasets, metrics, and pipelines to determine whether challenges are technically sound, reproducible, appropriately difficult, and genuinely require strong machine learning reasoning.

What You’ll Work On

You’ll review ML challenges involving:

  • Experiment design and model selection

  • Small and synthetic datasets

  • Data quality and preprocessing

  • Distribution shift and data contamination

  • Label noise and feature leakage

  • Model evaluation and metric selection

  • Hyperparameter tuning

  • Train / validation / test methodology

  • Reproducibility and deterministic pipelines

  • Statistical significance of model improvements

A key part of the role is determining whether a challenge actually rewards good ML reasoning, rather than simply being solvable through brute-force model selection or large hyperparameter searches.

What We’re Looking For

  • 3+ years of hands-on applied machine learning experience

  • Strong experience with:

    • ML experiment design

    • Model selection

    • Hyperparameter tuning

    • Model evaluation

    • Data preprocessing and validation

  • Strong understanding of train, validation, and test splits

  • Ability to identify:

    • Data leakage

    • Label noise

    • Distribution shift

    • Spurious correlations

    • Feature leakage

    • Data contamination

  • Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations

  • Strong understanding of ML evaluation metrics and when different metrics are appropriate

  • Experience debugging ML workloads across CPU and GPU environments

  • Ability to analyze technical problems and provide clear written feedback

Nice to Have

  • Experience creating or participating in Kaggle, DrivenData, or similar ML competitions

  • Experience designing benchmark datasets or ML challenges

  • Background in data-centric AI or dataset quality

  • Experience with synthetic data generation and validation

  • Familiarity with statistical testing, confidence intervals, and effect sizes

  • Experience with ML evaluation pipelines, RLHF, or AI model evaluation

  • Experience developing ML curricula or technical assessments

  • Understanding of common ML failure modes such as:

    • Shortcut learning

    • Spurious correlations

    • Goodhart’s Law

    • Simpson’s paradox

    • Metric gaming

What You’ll Be Responsible For

  • Reviewing ML challenges and determining whether they are well designed and technically solvable

  • Evaluating whether datasets contain meaningful and learnable signals

  • Identifying unintended shortcuts or artifacts in synthetic datasets

  • Determining whether tasks require genuine diagnosis of the underlying ML problem

  • Reviewing evaluation metrics and improvement thresholds

  • Detecting metric gaming, data leakage, and evaluation flaws

  • Verifying reproducibility across the complete data → model → evaluation pipeline

  • Assessing whether challenge difficulty is appropriately calibrated

  • Providing clear recommendations for improving, recalibrating, or excluding problematic tasks

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: Applied machine learning, experiment design, data quality, and model evaluation

This role is a strong fit for ML engineers who enjoy debugging experiments, understanding why models succeed or fail, identifying problems in datasets and evaluation pipelines, and designing rigorous machine learning experiments.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

AI Engineer - Colombia | English C1

Full Time Colombia $3500 - $4500 per month Software Development

Senior Software Engineer (Java Full Stack, Angular, React, Spring boot)

Other United States $111K - $150K per year Software Development

Senior Python Engineer (AI)

Full Time India Software Development

Software Engineer II

Full Time United States Software Development

Principal Application Architect

Full Time United States $118K - $169K per year Software Development

ERP Financial Consultant

Full Time United States $95000 - $130K per year Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Featuring 218,212+ Jobs in Software Engineer

Answer easy questions

Answer easy questions

218,212+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified