Job Title: Senior Software Engineer - Machine Learning
Location: 100% Remote (US Based Only)
- We cannot sponsor or transfer any visas, of any kind, at this time*
Hiring Manager: Senior Engineering Manager
Estimated salary range: $165,000 to $190,000
- The salary offered for this position will be based on a candidate’s experience and skill demonstrated during interviews and other evaluations
Job Description:
Ocient runs machine learning where the data lives. Models are trained and scored entirely inside the database through a native surface with no export, no separate Python stack, and no ETL round-trip. We support classification, regression, clustering, time series, neural network, ensemble, tree, pattern mining, and dimensionality reduction models today, and we are expanding that surface aggressively through 2026.
We are hiring a Senior Software Engineer to help build it. You will implement new model types, close feature and behavior gaps against the frameworks our customers already know, and make our ML run fast at scale. This is a hands-on engineering role on a small team with a lot of surface area, and you will own meaningful pieces of the roadmap end to end.
The work matters strategically. In-database ML is becoming table stakes across the industry, and our bet is on what comes next: in-database research, where optimization, spectral analysis, causal inference, simulation, and inverse modeling all become first-class declarative surfaces in the engine. The ML model catalog is the substrate that work stands on.
What You’ll Work On:
- Framework parity and behavioral correctness. Help close feature and semantic gaps against scikit-learn and Spark ML across the model catalog so our models behave the way users coming from those frameworks expect
- New model types and capabilities. Expand the model catalog across classification, regression, time series, and automated model selection, including loss functions and objectives we don't support today.
- Performance at scale. Make training and inference fast on datasets that don't fit anywhere else. Approximate nearest-neighbor acceleration, distributed optimizer tuning, and algorithmic work on the hot paths.
- Numerical and linear-algebra foundations. Help expand the SQL-level linear algebra surface (SVD, eigenvalues, matrix inverse and solve, sparse matrices) and spectral transforms, the substrate under PCS, regression, and optimization.
- Architecture. Our models are compiled into the query plan and execute as a native part of it, rather than running in a separate ML runtime. You'll work inside that architecture and help improve it.
- Collaboration and craft. Write clear design docs, tests, and documentation; investigate issues where behavior diverges from user expectations; and partner with Product, architects, and customer-facing teams to identify gaps before customers hit them.
Qualifications:
- 5+ years building production software systems, including solid experience in C++ (or comparable systems-level work in Java/Scala with a willingness to work primarily in C++).
- Hands-on experience implementing or integrating machine learning models in production. You have written the training loop, not just called into a library.
- Working knowledge of numerical methods: gradient-based optimization, loss functions and their gradients, numerical stability, feature scaling, convergence behavior.
- Familiarity with scikit-learn, Spark ML, XGBoost, or comparable frameworks, and awareness of where their defaults and semantics matter.
- Strong instincts around correctness, edge cases, and behavioral consistency – and the discipline to encode them in tests and documentation.
- Ability to work across teams and codebases and turn ambiguous requirements into concrete solutions.
An Exceptional Candidate Will Have:
- Experience comparing or validating model behavior across multiple ML frameworks.
- Experience with large-scale data systems, analytical databases, query planners, or distributed execution engines.
- Exposure to optimization (LP/QP/SOCP), spectral methods (FFT/DCT/DWT), causal inference, or probabilistic modeling for the in-database research surface we are building next.
- Experience with automatic differentiation or symbolic gradient generation.
- Familiarity with SQL internals like AST manipulation, expression rewriting, or planner integration.
What Success Looks Like:
- Customers see fewer surprises. Our models behave the way someone coming from scikit-learn or Spark ML expects, and where they differ, the difference is intentional and documented.
- Roadmap items such as ARIMA, quantile regression, AutoML, option parity, land with validated correctness against reference implementations.
- Feature gaps are identified from benchmarks and product analysis early, not discovered under customer pressure.
- You deliver across both parity work and broader ML initiatives, balancing short-term needs with long-term quality.