Please mention DailyRemote when applying
Volta builds and operates large-scale GPU compute infrastructure for AI workloads. Our platform is Kubernetes-native, spans multiple regions, and delivers virtual machines, storage, and networking through a fully automated infrastructure stack built on custom Kubernetes operators.
Platform Engineers work at the intersection of infrastructure and software development. You will translate three key inputs into durable platform capabilities: product roadmap requirements from the product team, operational learnings from the bring-up team, and security guidance from the security engineering team.
Design and implement Kubernetes operators and controllers that manage the lifecycle of compute, storage, and networking resources
Work closely with the product team to understand roadmap requirements and implement the platform capabilities that support them
Collaborate with the bring-up team to identify operational pain points and turn them into scalable platform features
Improve and extend the northbound API layer — the interface between user-facing services and the underlying infrastructure platform
Build and extend confidential computing capabilities across the platform stack — from secure bare metal and confidential VMs to Confidential Containers (CoCo)
Integrate security guidance from the security engineering team into platform-level controls and remediate security findings at the platform layer
Build platform capabilities around networking: reliability, performance, and observability of the overlay and underlay network stack
Contribute to storage platform improvements: provisioning workflows, attachment reliability, performance tuning, and failure handling
Own observability as a platform concern — instrument services, define meaningful metrics, and build tooling that gives the team visibility into platform health
Participate in code review, technical design discussions, and cross-team collaboration in an Agile (Kanban or Scrum) environment
3–5 years of software engineering experience, with a meaningful portion spent on infrastructure or platform systems
Working proficiency in at least one relevant language — Python, Go, or Rust — with experience writing production-grade backend services or automation, and a willingness to work across languages as the codebase evolves
Solid understanding of Kubernetes internals: the control loop model, CRDs, controllers/operators, and reliable reconciliation logic
Comfortable working close to the infrastructure layer — Linux, networking fundamentals, and distributed systems behaviour
Experience designing and building APIs or service interfaces that other teams depend on
Strong engineering fundamentals: clean code, testing, version control, code review, and CI/CD practices
Fluency with AI-assisted development, and interest in scaling agent-assisted workflows across the team (agentic CLI tools, MCP, skills, APIs) to amplify delivery.
Familiarity with confidential computing technologies: TEEs, AMD SEV, Intel TDX, or Confidential Containers (CoCo)
Experience integrating security requirements into platform or infrastructure systems
Familiarity with high-performance networking: overlay protocols, BGP, RDMA, or packet-processing frameworks
Hands-on experience with distributed storage systems (Ceph or similar) at an engineering level
Background building Kubernetes operators using frameworks such as Kopf, controller-runtime, or similar
Experience with observability tooling: Prometheus, Grafana, OpenTelemetry, or structured logging in distributed systems
Exposure to GPU infrastructure or HPC environments
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Platform Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!