Please mention DailyRemote when applying
Match your resume skills with our AI powered skill match!
Questions interviewers often ask for this role, with sample answers.
Upload your resume and we draft a letter for this exact role, tailored to what it asks for.
The role involves designing, building, and maintaining AWS infrastructure as code while managing production workloads and shared CI/CD tooling. You will also participate in an on-call support rota, handle security remediation, and automate manual operational tasks.
Company Profile
CloudMargin is an award winning, growth-stage FinTech company offering an innovative Software-as-a-Service (SaaS) service. We are rapidly challenging the way financial and corporate institutions manage the collateral and risk associated with derivatives trading. We believe in offering our community an affordable, easy to deploy and scalable service.
Backed by influential VC and corporate investors, our global team of 100 works closely across product, technology and sales disciplines. Our flat structure and ethos of openness and communication means you’ll be engaged with everyone in the business and have the senior leadership team as one of your key stakeholders.
Primary Purpose of Role
Reporting to the Lead SRE, you’ll be a senior hands-on engineer in the Platform team, a distributed team across the UK, Romania and Italy that owns the AWS estate behind CloudMargin’s SaaS platform, including the dedicated client environments we run for tier-one financial institutions. The role splits between building: advancing our multi-account AWS architecture, infrastructure as code and the shared tooling our development teams depend on; and running: taking a full share of the support rota, production incidents, security remediation and the compliance exercises our clients rely on. You will work closely with our development teams, Client Services and InfoSec, and you will be expected to automate away the manual work that currently takes up the team’s time rather than simply absorbing it.
Core Accountabilities
•
Design, build and maintain CloudMargin’s AWS infrastructure as code, using Terraform and Terragrunt across a multi-account estate and approximately 70 service repositories. •
Contribute to our AWS multi-account programme (AWS Organizations, IAM Identity Center, Service and Resource Control Policies, account bootstrap and baseline), including the per-client accounts that underpin our dedicated client environments. •
Build and maintain the shared CI/CD and release tooling our development teams consume: GitHub Actions reusable workflows, AWS CodeBuild and CodeDeploy, semantic-release, and policy-as-code checks such as Checkov. •
Operate containerised and serverless production workloads on AWS ECS Fargate and Lambda, together with the shared container registry and managed developer environments. •
Own monitoring, logging and alerting in Datadog: designing actionable alerts, reducing noise, and keeping observability spend proportionate to its value. •
Take a full share of the team’s weekly support rota and Datadog On-Call, responding to production incidents against the team’s internal SLOs and producing clear root cause analysis. •
Operate the platform’s security tooling (Wazuh, Prowler and Trivy) and drive vulnerability findings through to remediation within agreed timescales. •
Manage the perimeter and edge: AWS WAF, CloudFront, firewalls, client IP whitelisting, and the full certificate lifecycle across public, internal and client-facing endpoints. •
Support CloudMargin’s compliance obligations across ISO 27001, DORA and SWIFT, including the annual penetration test and the annual disaster recovery exercise, which is run with client participation. •
Operate and upgrade the data and messaging platform: Aurora MySQL, RDS Proxy, DocumentDB, Redis and Amazon MQ, planning version and architecture migrations around agreed change windows. •
Deliver client-facing infrastructure onboarding (SFTP and AWS Transfer Family connections, key management, tenant and sandbox provisioning), working directly with Client Services and, where needed, with client technical teams. •
Manage secrets, identity and access across AWS Secrets Manager, SSM Parameter Store, IAM Identity Center, Active Directory and Tailscale. •
Contribute to cost optimisation (FinOps): cost and usage analysis, budgets and guardrails, rightsizing, and backup and snapshot lifecycle management. •
Relentlessly automate toil, replacing manual runbook steps with self-service tooling so that the team’s capacity goes into engineering rather than repetition. •
Modernise and decommission legacy systems safely and incrementally, without disrupting a platform clients depend on every working day. •
Reduce complexity and avoid over-engineering, while ensuring the considerations of performance, scalability, security and testability are addressed when designing and refining solutions. •
Document your work in runbooks, architecture decision records (ADRs) and threat models, so that technical specifications, system designs and project details are accessible to both technical and non-technical stakeholders. •
Work to Agile practices: two-week sprints, Definition of Ready and Definition of Done, backlog refinement, and honest reporting of risks, progress and milestones to stakeholders. •
Share knowledge and support the development of less experienced engineers across a distributed, multi-site team.
Experience • At least 5 years of practical experience in Platform Engineering, DevOps, SRE or Infrastructure roles, or equivalent. •
A minimum of 3 years of hands-on experience designing and running production workloads on AWS. •
At least 2 years managing production environments on AWS ECS, Fargate and serverless (Lambda); Kubernetes experience is transferable, but note that our estate runs on ECS Fargate rather than Kubernetes. •
Demonstrable experience writing and maintaining production Terraform at scale. Terragrunt, or an equivalent orchestration layer for multi-account and multi-environment deployments, is highly desirable. •
Experience of AWS multi-account architecture (AWS Organizations, IAM Identity Center, Service Control Policies and account baselining) is highly desirable. •
Experience of carrying production on-call and leading incident response in a customer-facing environment. •
Hands-on experience of vulnerability management and remediation, and of working within a recognised security or regulatory framework such as ISO 27001, SOC 2 or DORA. •
Comfortable reading and debugging Node.js services, and able to write production-quality Python and Bash. •
Prior experience within a financial institution, or in a regulated SaaS environment serving financial services clients, would be highly desirable.
Skills • Depth and breadth across AWS: ECS/Fargate, Lambda, EC2, RDS and Aurora MySQL, DocumentDB, ElastiCache/Redis, Amazon MQ, S3, EFS, ELB/ALB, Route 53, API Gateway, CloudFront, WAF, ACM, Secrets Manager, SSM Parameter Store, KMS, AWS Backup, ECR, CodeBuild, AWS Transfer Family, Organizations and IAM Identity Center. •
Expertise in Terraform for infrastructure automation, including module design, remote state and workspace strategy; working knowledge of Terragrunt. • Skilled with Ansible for configuration management and application deployment of the remaining EC2-based estate. •
Proficiency in implementing CI/CD pipelines with GitHub Actions and AWS CodeBuild/CodeDeploy, including automated versioning and release (semantic-release or similar). •
Familiarity with observability in Datadog: metrics, logs, APM, database monitoring and on-call scheduling. Experience of migrating from legacy tooling such as Nagios or the ELK stack is useful. •
Security tooling: SIEM and file integrity monitoring (Wazuh or similar), cloud security posture management (Prowler, AWS Security Hub), container and dependency scanning (Trivy), and policy-as-code (Checkov). •
A deep understanding of networking and system architecture: VPC design, peering and endpoints, DNS, TLS and certificate management, zero-trust remote access (Tailscale or similar) and firewall administration. •
Linux administration on Ubuntu (CentOS experience is useful for migration work), plus knowledge of both Nginx and Apache web servers. •
Proficiency in scripting and automation with Python and Bash, and comfort working in HCL, YAML and SQL. •
Solid understanding of microservice and event-driven architecture, including SQS, SNS and EventBridge, and of RESTful APIs and Single Page Applications (SPA). •
An understanding of FinOps practice: cost allocation and visibility, budgets and guardrails, rightsizing and unit economics. •
Experience with Atlassian tools such as Confluence for documentation and Jira for project tracking, and openness to AI-assisted engineering tooling, which the team uses actively. •
Knowledge and application of Agile principles and values as outlined in the Agile Manifesto. •
A pragmatic and proactive approach to problem-solving, with a strong bias towards automating repetitive work. •
Ability to positively handle and thrive under pressure, including during production incidents and planned weekend change windows. •
A strong team player who prioritises collective team success over individual achievements. •
Excellent written and verbal communication, and the ability to work cross-functionally across time zones with development teams, Client Services, InfoSec and client technical staff. •
Highly motivated, collaborative, and eager to contribute to a positive work environment.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Featuring 217,169+ Jobs in Platform Engineer
Answer easy questions
217,169+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!”