Job Summary
This role has two equal parts. The first is building and maintaining the data infrastructure that makes Ralco's Innovation pipeline trustworthy and queryable: designing pipelines, enforcing data quality, and doing the real, hands-on work of entering, structuring, and validating data as it comes in from the lab, the barns, and the field. This includes building, or working with IT to build, a LIMS-style direct-entry system so that work happens correctly at the source rather than through repeated manual re-entry.
The second is working alongside members of the R&D team and IT to develop models and tools that surface candidate compounds and patterns the science team can act on. This includes actively reviewing data across projects and bringing forward insights, trends, or connections for the science team to think through and possibly act on, while safeguarding the accuracy and confidentiality of Ralco's most valuable data assets. Success means the Innovation team spends less time hunting for or correcting data and more time doing science, with a shared database that grows more useful, trustworthy, and queryable over time rather than more unwieldy, and this person is known across the team both as someone who keeps the data itself in order and as someone who consistently adds something useful to the team's work.
This job description is not meant to be all-inclusive, and duties may vary based on departmental and company needs.
Key Responsibilities
- Design, build, and continuously improve the pipeline connecting lab instruments, in-vitro work, and barn/field data sources into the shared database
- Build, or work with IT to build, a LIMS-style direct-entry system, owning the correct entry, tagging, organizing, and storing of it correctly in the shared database
- Design and maintain the database schema and tagging conventions so the data can be queried and modeled correctly
- Work with IT to design and validate permission, gated access controls that determine who can see sensitive data, testing them before any sensitive data goes live
- Ensure new compound and trial data is entered with a complete, validated profile
- Resolving data-quality issues directly with the originating scientist before any correction is made, and lead periodic data audits to reconcile duplicate or conflicting records and standardize naming conventions
- Build and maintain anomaly detection, compound similarity/clustering, and pattern-flagging tools that are useful well before a full predictive model is trustworthy
- Build, retrain, and validate the models used for querying the database on a documented cadence tied to data volume, tracking hit-rate before and after each cycle and documenting the trend
- Work with the R&D team to make sure the model used for querying stays the most accurate and correct fit for the data, revisiting its underlying design when it no longer fits rather than assuming it always will
- Actively review data across active research projects on an ongoing basis, bringing forward insights, trends, or connections for the science team to think through and possibly act on
- Conduct regular cross-project data reviews to help surface new angles or connections for the R&D team on active research tracks
- Run patent searches and compile literature searches on request to support hypothesis-building, manuscript preparation, and Regulatory's IP review, maintaining a running log so the same ground is never covered twice
- Produce a quarterly gap-audit report, and a brief report after every model retraining cycle documenting what changed and what the new hit-rate is
- Follow, and help enforce, Ralco's IP protection practices for anything touching the shared database, escalating promptly any data handling or access issue that could put proprietary information at risk
- Regularly meet with the R&D and IT teams to stay closely connected to their work, translating between researcher needs and data architecture as new data types and sources come online
Key Competencies
- Data pipeline design and hands-on database schema architecture, including LIMS-style direct-entry system design or integration
- SQL and at least one scripting language for pipeline and data work; experience with permissioned/gated data-access systems, familiarity with MCP architecture or similar data governance frameworks a plus
- Practical, hands-on machine learning experience, model training, validation, retraining, and performance evaluation, including anomaly detection, clustering, and similarity algorithms on structured datasets
- Sound judgment handling proprietary and sensitive data
- Self-directed learning; demonstrated ability to track how AI/ML tools and techniques evolve and evaluate new approaches on one's own initiative
- Comfortable working across scientific and technical teams; able to translate between researcher needs and data architecture
- Ralco Values and Purpose Alignment
- Values: Courageous Curiosity, Driven to Deliver, Master Your Craft, Root for the Team, Earn the Relationship
- Purpose: Ralco creates effective natural solutions that address agriculture's greatest challenges.
Qualifications
Education: Bachelor's or Master's degree in Data Science, Data Engineering, Bioinformatics, Computer Science, or a closely related field
Experience: 3–6 years of experience in data science, data engineering, bioinformatics, master data management, or a closely related field
Technical: Real experience designing data pipelines; SQL and relational database concepts; practical, hands-on machine learning experience; experience with permissioned/gated data-access systems.
Travel: Regular travel to Marshall, MN expected.
Location: Remote (US)
Preferred Qualifications:
- Experience building or integrating a LIMS-style direct-entry system
- Direct, hands-on experience building a predictive model end-to-end: preparing training data, training and validating a model, and deploying and retraining it as new data comes in, not just coursework or theoretical exposure
- Experience designing a database schema and tagging system from scratch for messy, real-world data, deciding how to structure and label information so it stays queryable and useful as it grows, not just working within a schema someone else already built
- Demonstrated self-directed learning; concrete examples of teaching oneself a new technical or scientific domain without formal instruction
- Shows curiosity about the underlying science, not just the data structure; asks why data looks the way it does
- Some background or coursework in life sciences, chemistry, or animal/agricultural science
- Experience running patent database searches
About Ralco
Ralco is a family-owned agricultural health and nutrition company with a global presence, distributing products in over 40 countries. Inspired by nature and powered by science, Ralco develops innovative, natural solutions to support plant and animal health.