Please mention DailyRemote when applying
See how much of this job your resume covers, and what’s missing.
Want a recruiter to go through it line by line?
Get professional reviewUpload your resume and we draft a letter for this exact role, tailored to what it asks for.
You will bridge the gap between applied machine learning and platform engineering by building reliable, observable, and production-ready ML workflows. Responsibilities include refactoring models, automating evaluation pipelines, and implementing robust monitoring for performance and cost.
About Smart Working
At Smart Working, we believe your job should not only look right on paper but also feel right every day. This isn’t just another remote opportunity — it’s about finding where you truly belong, no matter where you are. From day one, you’re welcomed into a genuine community that values your growth and well-being.
Our mission is simple: to break down geographic barriers and connect skilled professionals with outstanding global teams and products for full-time, long-term roles. We help you discover meaningful work with teams that invest in your success, where you’re empowered to grow personally and professionally.
Join one of the highest-rated workplaces on Glassdoor and experience what it means to thrive in a truly remote-first world.
About the Role
As a Senior ML Engineer, you will bridge the gap between applied machine learning and robust platform engineering. You will focus on translating AI capabilities into reliable and observable system components.
You will help ensure that machine learning models meet strict quality, cost and latency thresholds before reaching production. The role spans production ML architecture, workflow reliability, model evaluation, provenance, testing, human feedback integration and AI provider observability.
\nRefactor and upgrade existing ML models, including NLP, generative AI and transcription models, into standardised, production-ready modular contracts.
Engineer resilient ML workflows using DAG-based orchestration tools such as Argo Workflows and Airflow.
Ensure robust retry logic, error handling and repeatable execution across ML workflows.
Define, automate and maintain strict ML evaluation pipelines using golden datasets.
Ensure new models meet baseline quality and performance thresholds before release.
Design and implement tracking mechanisms to capture precise model, prompt and input data provenance for full auditability and reproducibility.
Build infrastructure for shadow testing and A/B testing on live traffic.
Implement reliable fallback and kill-switch mechanisms to support safe deployment.
Engineer structured feedback pipelines that capture human reviews and corrections to continuously enrich and refine training datasets.
Integrate third-party AI APIs and manage adapter interfaces.
Implement granular telemetry to track compute costs, token usage and latency.
Proven track record operating at the intersection of Machine Learning, MLOps and Platform/Backend Engineering.
Deep understanding of evaluation metrics for generative AI, LLMs and/or speech models, including Precision, Recall, F1, WER/CER, groundedness and hallucination rates.
Strong hands-on experience with containerisation using Docker and Kubernetes, alongside modern orchestration frameworks.
Experience building observable ML systems with logging, monitoring and strict cost-attribution tracking.
Familiarity with data provenance, compliance standards and building fail-safe mechanisms for sensitive data and AI outputs.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 216,193+ Jobs in Software Development
Answer easy questions
216,193+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”