The role involves leading the cloud infrastructure and DevOps function, including designing, deploying, and securing product environments on Google Cloud. The engineer will act as a player-coach, mentoring team members while remaining hands-on with infrastructure automation and incident response.
Lead DevOps & Cloud Infrastructure Engineer
What this Job Entails
This is a hands-on lead role owning the cloud infrastructure and DevOps function for Astreya's product portfolio. The person in this seat is accountable for how our products get built, deployed, secured and run — in our own environments and inside customer environments.
Google Cloud is the primary platform; Azure is secondary. This is not a role that sets direction and hands the work to someone else: the expectation is that this person writes the Terraform, debugs the cluster, runs the deployment, and then mentors the engineers who will do it next time. They act as the subject matter expert for infrastructure across product, engineering and delivery teams, and as the technical owner in customer conversations about deployment, security and hosting models.
Scope
Owns the infrastructure and deployment layer end to end, from CI pipeline to production runtime, across multiple products on independent release cadences
Resolves complex problems where the diagnosis requires in-depth evaluation of many interacting variables — networking, identity, cluster behaviour, managed service limits, cost
Exercises independent judgment in selecting platform services, patterns and tooling, and is expected to defend those choices technically and commercially
Operates as a player-coach: a small team reports to this role, and the role remains individually productive
Own network and perimeter design: VPC architecture, Private Service Connect, Cloud NAT, Cloud Armor, private connectivity to customer systems
Own secrets, keys and identity: Secret Manager, Workload Identity, IAM design and least-privilege service accounts, KMS, certificate lifecycle
Run R&D on platform services we haven't used yet — evaluate the GCP service, identify the Azure and AWS equivalents, and produce a recommendation with a working proof of concept, not a slide
Manage cloud cost: attribution by product and environment, budget alerts, rightsizing, commitment planning
Deployment into environments (a core reason this role exists)
Make deploying our products into a new environment a repeatable, documented, low-drama exercise — both customer-VPC and Astreya-hosted models
Own the deployment artefacts: Helm charts, namespace and pod topology, environment configuration, migration and rollback paths
Work directly with customer infrastructure and security teams during onboarding — answer their architecture and security questions, adapt to their constraints, and get us to production
Produce and maintain the evidence infrastructure security reviews and audits ask for: architecture diagrams, data flow, access controls, encryption posture, audit logging, retention
DevOps and platform engineering
Own CI/CD across products: build pipelines, artefact promotion, environment strategy, source control branching strategies across multiple release cadences
Write automation and internal tooling for provisioning and operating infrastructure — infrastructure as code (Terraform) as the default, scripts and services where it isn't
Build the observability layer: metrics, logs, traces, alerting, dashboards and SLOs that let us find problems before customers do
Improve scalability, reliability, capacity and performance of the platform, including the infrastructure serving AI/LLM workloads
Harden the platform: image scanning, dependency and vulnerability management, network policy, secrets hygiene, patching
Leadership and process
Lead and mentor DevOps and infrastructure engineers; set standards and review their work
Work with product owners and engineering leads to understand requirements, surface infrastructure bottlenecks early, and propose resolutions
Own incident response for platform issues and produce root cause analyses for outages that are honest about cause and specific about prevention
Design and implement improvements to existing support processes and tooling; introduce innovations and follow through on execution, not just proposal
Document decisions, runbooks and resolution history so the next person doesn't rediscover them
Other duties as required. This list is not a comprehensive inventory of all responsibilities assigned to this position
Required Qualifications/Skills
Bachelor's degree (B.S/B.A) from a four-year college or university and 8+ years' related experience and/or training; or an equivalent combination of education and experience
Deep, hands-on Google Cloud experience — has designed and run production workloads on GCP, not just passed a certification
Strong Kubernetes experience in production: GKE preferred, including networking, ingress, autoscaling, resource management and debugging failures under load
Infrastructure as code at a professional standard — Terraform, module design, state management, review discipline
Strong scripting and coding ability (Python, Go, Bash or equivalent) and enough software development experience to work inside application repositories, not just around them
CI/CD ownership across multiple products and release trains, with source control branching strategies
Monitoring, alerting and incident management experience, including writing the RCA afterwards
Experience with open-source tooling in large distributed systems
Good communication skills — can hold a technical conversation with a customer's security team and a working conversation with a developer on the same day
Ability to pick up unfamiliar technology quickly and carry several threads at once
Preferred Qualifications
Working Azure knowledge (AKS, Key Vault, Entra ID, VNet) and the ability to translate a GCP design onto it; AWS exposure a bonus
Experience deploying a product into customer-controlled cloud environments as a vendor
Experience supporting AI/ML workloads in production — Vertex AI, model endpoints, GPU scheduling, inference cost control
Exposure to SOC 2, ISO 27001 or comparable audits from the infrastructure side
Experience integrating with enterprise ITSM platforms (ServiceNow) and the connectivity that requires
Prior experience as a first or early infrastructure hire on a product team
Relevant certification (Google Professional Cloud Architect, Professional Cloud DevOps Engineer, CKA)
Physical Demand & Work Environment
Must have the ability to perform office-related tasks which may include prolonged sitting or standing
Must have the ability to move from place to place within an office environment
Must be able to use a computer
Must have the ability to communicate effectively
Some positions may require occasional repetitive motion or movements of the wrists, hands, and/or fingers
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”