Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
You will own the trust layer of an AI agent for industrial plants, focusing on reliability, grounding, and performance. This involves diagnosing agent misbehavior, building evaluation harnesses, and implementing mechanisms to ensure output quality and accuracy.
We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
Project - the aim you'll have
Our client builds an AI copilot for process engineers in oil refineries and chemical plants: a natural-language interface where engineers ask questions about live plant data — equipment, sensor tags, process trends — and get grounded, chart-backed answers. The users are experienced engineers who are rightly skeptical of AI: in this domain, a fabricated number or a silent wrong assumption has real cost. The product wins or loses on whether the agent can be trusted.
This role owns the trust layer of that agent inside a large, active Python codebase. It is not feature work with an LLM endpoint bolted on. The work is the mechanics of agent reliability: making the agent say "I don't know" instead of inventing, surfacing every assumption it makes so the user can correct it, grounding every claim in actual data, holding output quality through model migrations, and keeping latency acceptable while doing all of the above.
To make the day-to-day concrete, this is what the engineer currently in this seat shipped in the last four months (all of it flag-gated, in small PRs, reviewed async daily by a team spread across the US and Australia):
- An assumption auditor: detects the silent assumptions the agent makes when answering (which equipment, which time window), validates them via multi-draw consensus, and surfaces them in the UI as correctable chips — the engineer can fix an assumption and rerun the analysis.
- A grounding auditor that catches reports fabricated from empty data feeds before they reach the user.
- An adversarial reviewer sidecar that critiques generated charts for correctness before display.
- Successive frontier-model evaluations (loop behavior, directive adherence, regression on a replay harness) that decided when to flip the product's default model — including, twice, deciding NOT to flip.
- Hardening of a plant-exploration tool against hallucinating structure that the data does not support.
- A latency fix: a narration side-loop was inflating query response times; capped it and made it best-effort.
- A concurrency fix making a shared data-reset path atomic, eliminating intermittent production read errors.
If reading that list is more interesting to you than building another CRUD feature, this role is for you.
Expectations - the experience you need
Nice to have
What you will do
Our Benefits
We are accepting applications from LATAM countries
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”