Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
The Biomedical SME will map dbGaP variables to standardized vocabularies and design expert-adjudicated reference datasets for AI challenge evaluation. They will also provide technical authority on scientific integrity, benchmark validation, and FHIR-based interoperability.
About Mind Moves
Mind Moves is a women-owned Washington, D.C.-based firm that helps government and business partners navigate digital transformation using "human-in-the-loop" AI. Our expert team has delivered responsibly developed AI products that drive millions in impact across agencies like the National Institutes of Health (NIH).
Program Overview
The National Library of Medicine (NLM) is launching Biomedical AI Challenges — a structured prize competition program designed to catalyze innovation in AI-powered tools for biomedical research infrastructure. The program seeks to improve AI-enhanced semantic search across PubMed/PMC and establish foundational data interoperability by mapping raw, cohort-specific dbGaP study variables directly to precise, standardized international vocabularies/ontologies (such as LOINC, RxNorm, SNOMED CT, and the UMLS Metathesaurus).
Why This Work Matters
Right now, dbGaP study variables and metadata are described in disparate, cohort-specific shorthand rather than standardized codes. Because standardized terminologies attach stable, unambiguous codes to clinical concepts, they are what make it possible for different studies, systems, and AI tools to "talk" about the same phenotype, lab result, or exposure in the same way. Without that shared vocabulary layer, cross-study comparison stays manual and error-prone, dbGaP's rich phenotypic data remains difficult to discover, and researchers can spend months pursuing a controlled-access request only to find the cohort doesn't match their needs. The SME's terminology mapping and curation work is the foundation this entire Challenge is built on: it is what turns free-text variable descriptions into the standardized, computable concepts that both Challenge tracks — and, ultimately, the broader research community — depend on for reliable semantic search and cross-study interoperability.
Position Summary
The Biomedical Subject Matter Expert (SME) serves as a senior scientific advisor and technical authority to provide expert guidance on evaluation framework development, benchmark dataset validation, and scientific review. The Biomedical SME contributes meaningfully to shaping the scientific integrity and rigor of challenge-related tasks.
Key Responsibilities
Map dbGaP variables to controlled vocabularies and common data elements, including PhenX measurement protocols, using the dbGaP data dictionary's VARIABLE_SOURCE and SOURCE_VARIABLE_ID fields, and classify mappings by confidence level (e.g., identical, comparable, or related) consistent with dbGaP/PhenX conventions.
Curate and validate semantic annotations linking dbGaP variables to standard terminologies such as the UMLS Metathesaurus, LOINC, and UMLS Concept Unique Identifiers (CUIs), following the approach used by groups like NLM's Lister Hill Center, Medical Data Models, and NHLBI's TOPMed program.
Support FHIR-based interoperability by ensuring annotated vocabularies are correctly represented in the dbGaP FHIR schema and accessible via the dbGaP FHIR API.
Reference Set & Evaluation Design
Design and curate the expert-adjudicated reference ("gold standard") mapping set against which Track 1 participant submissions are scored, including explicit criteria for the ONT_NONE discard class (variables with no valid target-vocabulary concept).
Author and maintain annotation guidelines and adjudication protocols so reference-set decisions are reproducible and defensible under review or challenge.
Run or oversee inter-annotator agreement checks across SME reviewers and resolve disagreements before a mapping enters the reference set.
Advise on Track 2 relevance judgments — labeling dbGaP studies as relevant/not relevant to natural-language research queries — in a form suitable for computing ranking metrics.
Contribute to held-out or adversarial "probe" cases used to detect memorized or hardcoded submissions during anti-gaming review.
Required Qualifications
Advanced degree (MS or PhD preferred) in bioinformatics, biomedical informatics, epidemiology, genetics, library/information science, or a related field — or equivalent professional experience.
Demonstrated experience with biomedical controlled vocabularies and terminology standards (e.g., UMLS, LOINC, MeSH, PhenX, or OBO Foundry ontologies such as HPO or OBI).
Familiarity with clinical/genomic data repositories and data-sharing models comparable to dbGaP (e.g., NIH CDE Repository, TOPMed, other GWAS/genomic consortia).
Experience with semantic web or ontology tooling (e.g., OWL, RDF, Protégé).
Experience with Python or R.
Strong written communication skills for documenting mapping rationale.
Working knowledge of information-retrieval and classification evaluation metrics used to grade automated systems (precision, recall, F1-score, NDCG) and how to translate expert judgment into a defensible ground-truth/reference dataset.
Preferred Qualifications
Prior experience working directly with dbGaP data dictionaries, submission packets, or the dbGaP FHIR API.
Experience with data-dictionary harmonization tools (e.g., D2Refine or similar).
Familiarity with GWAS/genomic consortium data harmonization processes.
Experience with FHIR resource modeling in a research-data or genomics context.
Prior experience building or contributing to benchmark/gold-standard datasets for an evaluation campaign or shared task (e.g., TREC-style relevance judgments, NLP/IR bake-offs) and calibrating multiple annotators to a shared rubric.
Position Type: Part-Time Contract (Independent Contractor)
Program: NLM Biomedical AI Challenges Program
Period of Performance: Sept 2026 – March 2027
Project Hours: 500 – 600 total hours
Compensation: $130-160/hr
Location: Remote within the US
Interview Process
Priority deadline for application is September 18th, 2026, and we look to fill these roles as soon as possible. If interested, please submit your application today! We expect the process to include a short hiring exercise followed by 1 20-30 min interview for final candidates.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Teaching
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”