For Employers

Bitdeer Technologies Group

Staff MaaS Backend Engineer

Posted 11 days ago
$180K - $260K per year
10+ years experience
Apply Now

Please mention DailyRemote when applying

?/100
Resume Match Score

Match your resume skills with our AI powered skill match!

Get professional review
AI Summary

The Staff Backend Engineer will co-own the end-to-end design and implementation of a globally distributed, multi-tenant Model-as-a-Service platform. This role involves driving technical strategy, optimizing performance for high-throughput AI inference, and ensuring robust reliability and billing accuracy.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit https://ir.bitdeer.com/

Position Overview

We are seeking a Staff Backend Engineer to take our Model-as-a-Service (MaaS) working system and re-architect it into a commercial, globally distributed, multi-tenant token service — one that sustains at least millions of monthly active users, high sustained token throughput per GPU, and invoice-grade accounting, with scalability, reliability, and observability engineered in deliberately rather than absorbed under load. This role works directly with the principal architect: co-owning the MaaS system design as the deepest backend voice in that conversation, and owning the implementation end to end — the code that ships, the migrations that land, and the service that stays up. Design authority is shared; delivery accountability is not. The mandate is explicit: measure what exists, find where it breaks before it breaks in front of a paying customer, and carry the platform there incrementally — with each step independently shippable, reversible, and non-disruptive to the tenants already on it. The role is deeply hands-on: it reads and rewrites the existing Go services, owns SLOs and on-call for a revenue-bearing service, and sets the backend engineering standard for the MaaS team.

Key Responsibilities

  • Architecture, Strategy, and Leadership: Co-own the end-to-end MaaS system design with the Principal Architect, authoring decision records and defending technical trade-offs. Drive the platform through its maturity roadmap by delivering operable, measurable capabilities rather than mere demos. Lead technical execution by setting stringent Go and API standards, mentoring engineers, and aligning cross-functional teams.
  • Inference Gateway and API Surface: Own the wire compatibility contract for major formats (OpenAI, Anthropic), supporting advanced features like streaming, tool calling, and structured output. Evolve the routing tier to handle load-aware, model-aware, and prefix-cache-aware endpoint selection with robust circuit breaking and fallback mechanisms. Run versioning and deprecation as a published contract to guarantee external customer code stability across underlying changes.
  • Performance Optimization and Model Lifecycle: Maximize platform economics and performance by optimizing token throughput, KV cache tiering, and time-to-first-token (TTFT) latency at the p95/p99 levels. Mature the model serving control plane by integrating deployment tooling, LoRA multiplexing, and cold-start-aware autoscaling directly with the Kubernetes fleet. Treat regressions in cost-per-million-tokens or latency metrics as critical system incidents.
  • Global Topology and Reliability (SLOs): Scale the platform to a globally distributed architecture featuring regional inference pools, capacity-aware failovers, and an active-active control plane. Define, publish, and rigorously defend strict Service Level Objectives (SLOs) baselined against actual system performance rather than aspirations. Ensure operational resilience through peak-concurrency load testing, robust on-call runbooks, and predictable load-shedding during overloads.
  • Security, Identity, and Multi-Tenant Isolation: Enforce fail-closed authorization, robust multi-tenant isolation, and zero-retention data paths across the network, cache, and storage layers. Manage the complete lifecycle of API keys and OAuth credentials while distributedly enforcing rate limits and quotas without relying on client-supplied identifiers. Design strict abuse, rate, and prompt-injection controls, treating all user and model-generated content as untrusted data.
  • Token Metering and Billing Cor
  • rectness: Build an idempotent, exactly-once metering system to capture uncached, cached, output, and reasoning tokens accurately across all requests. Maintain a high-volume usage ledger that enforces prepaid spend caps asynchronously and reconciles perfectly with the invoicing system. Safeguard business integrity by treating any metering or billing defect as a critical revenue and trust incident.
  • Safe Migrations and Observability: Execute zero-regression, incremental platform upgrades using strangler-style replacements, shadow traffic, and stateful dual-writes. Deliver end-to-end request tracing and cost telemetry, defining internal schemas for token and quality attributes. Enforce absolute log hygiene by ensuring prompts, completions, and PII are never retained outside of explicit, consented policies.

Qualifications

  • Bring 8+ years of backend engineering experience, including 3+ years owning a high-traffic, multi-tenant API platform for paying customers. Act as a senior technical voice who can collaborate closely with architects, write rigorous design documents, and commit to executing architectural decisions effectively.
  • Demonstrate expertise in scaling distributed systems through multi-region active-active deployments, caching, backpressure, and targeted performance engineering that measurably lowers unit costs. You must have a proven track record of safely executing zero-downtime brownfield migrations for stateful subsystems—like metering or ledgers—without regressions or accounting gaps.
  • Possess deep hands-on proficiency with Go-based services, production Kubernetes (including Envoy and GPU-aware scheduling), and the architectural trade-offs of datastores like PostgreSQL, Redis, and Kafka. Additionally, you will drive operational visibility by owning end-to-end observability strategies using OpenTelemetry and high-cardinality analytics stores.
  • Apply a systems-level understanding of LLM serving to manage complexities like server-sent-event streaming, KV/prefix caching, and the trade-offs between time-to-first-token (TTFT) and throughput. Leverage this foundation to build highly reliable, exactly-once metering and billing systems that accurately reconcile billions of events under partial failure conditions.
  • Enforce strict multi-tenant security disciplines by designing fail-closed authorization, mandating verified identities, and guaranteeing absolute cross-tenant isolation. Bring operational maturity to a revenue-bearing platform by carrying on-call responsibilities, running blameless incident reviews, and translating outages into structural improvements.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Automatically Apply to the Best Remote Jobs

Stop the endless job search. Our AI finds and applies to the best jobs for you.

Try it Now
Keep looking

Similar Jobs

See all Remote Software Development jobs →

Bridge Project Engineer

Full Time United States Software Development

Senior Virtualization Engineer - Managed Services

Full Time United States $106K - $150K per year Software Development

e-Learning Developer (Remote Contract Opportunity)

Freelance Philippines Software Development

Senior Software Engineer

Full Time Brazil Software Development

Brazil- Remote: Full Stack AI Developer

Full Time Brazil Software Development

Backend Engineer

Full Time Worldwide 50000 - 80000 per month Software Development
Apply Now

Personalize your Remote Job Search in 3 Easy Steps!

Discover remote opportunities in Backend Engineer

Answer easy questions

Answer easy questions

200,000+ jobs across 15+ categories

Get your best job matches

Get your best job matches

Only hand-screened, legit jobs

Find a remote job faster

Find a remote job faster

No ads, scams, or junk

I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”

Sarah J. — Sarah J. · Marketing Manager ★★★★★ Verified