Please mention DailyRemote when applying
About the team:
Platform Infrastructure builds, operates, and continuously evolves FYUL's container platform and cloud foundation. We foster a DevOps culture through self-service tooling, enabling product engineering teams to ship reliable, secure, and cost-efficient services as the business scales. The team owns our AWS cloud accounts, Kubernetes platform, cloud networking, observability stack, core databases, CI/CD pipelines, and infrastructure-as-code, and acts as the go-to partner for engineering teams on cloud and DevOps topics.
About the role:
We're hiring a Senior SRE II to join Platform Infrastructure as one of the team's senior individual contributors. At this level, you're the go-to person for our most complex infrastructure problems: you architect and drive large-scale automation and reliability initiatives, set standards other engineers follow, and mentor Associate and mid-level SREs. You'll split your time between hands-on platform work - Kubernetes, AWS, GCP, CI/CD, observability - and technical leadership: proposing designs, reviewing others' work, and helping the team make good build-vs-buy and cost/reliability trade-offs.
Your daily tasks will include:
Infrastructure & reliability: Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts and environments using infrastructure as code.
Design and operate our Amazon EKS clusters, including networking policies, persistent storage, and scaling strategies for containerized workloads.
Own and evolve core platform services: cloud networking, Kubernetes, and the databases and messaging systems engineering teams depend on.
Automation & infrastructure as code: Drive large-scale automation projects and set standards for using Terraform / Terragrunt and GitOps (ArgoCD) across teams.
Lead adoption of automation to reduce manual operational work and keep environments consistent and repeatable.
Observability & incident response: Be the go-to person for solving complex, cross-service infrastructure problems.
Drive initiatives that improve reliability and observability (Grafana, Prometheus, Loki, Tempo, Mimir) so systems scale with minimal manual intervention.
Participate in on-call rotation, lead incident response for production issues, and write clear runbooks, ADRs, and postmortems.
Security & cost efficiency: Lead security efforts within the team - IAM, encryption, secure logging - and mentor others on secure infrastructure practices.
Audit infrastructure spend regularly and drive cost optimization across the platform (rightsizing, autoscaling, FinOps practices).
Collaboration & mentorship: Mentor mid-level SREs, provide detailed feedback, and support onboarding of new team members.
Communicate complex technical concepts clearly to both engineers and non-technical stakeholders.
Partner with product engineering squads to understand their needs and represent Platform Infrastructure in cross-team initiatives.
Your qualifications:
These reflect the technical bar we hold Senior SRE II's to internally, based on our SRE competency framework and current stack.
Core technical experience:
Solid Linux systems administration background and comfort scripting in Python.
Strong AWS knowledge: EKS, IAM (roles, policies, IRSA), VPC networking, RDS, S3, SQS, and familiarity with the Well-Architected Framework; experience in multi-account AWS environments is a strong plus.
Hands-on experience operating and troubleshooting Kubernetes (EKS) at production scale, including Helm chart development, CNI networking (we run Cilium), pod networking/IPAM concepts, and container security (ECR, image scanning).
Proficiency with Terraform (modules, state management) and ideally Terragrunt for multi-environment management; GitOps experience with ArgoCD.
Experience with Postgres, MySQL and/or MongoDB in production scale, including Aurora.
CI/CD experience with Jenkins (Jenkinsfile, shared libraries) and/or GitHub Actions, and familiarity with deployment strategies such as blue-green and canary.
Experience with the Grafana observability stack (Grafana, Prometheus, Loki, Tempo, Mimir) - metrics design, dashboarding, alerting, log aggregation, and distributed tracing. Not only using but also maintaining it.
Practical incident management experience: on-call rotations, structured incident response, and writing runbooks/postmortems.
Working knowledge of 12-Factor App principles and cost optimization / FinOps awareness.
2. How you work:
A methodical, data-driven approach to troubleshooting rather than guessing.
Strong written communication - you write runbooks, ADRs, and postmortems that others can actually follow.
Comfortable driving initiatives with ambiguous ownership, and taking accountability for outcomes rather than waiting to be asked.
Track record of mentoring less senior engineers and giving direct, constructive feedback.
Several years of hands-on production infrastructure/SRE experience, with demonstrated ownership of initiatives at a senior individual-contributor level (leading design work, setting standards, being the escalation point for hard problems).
3. Nice to have:
GCP Experience.
Experience with Kafka / AWS MSK.
Prior experience in regulated or compliance-sensitive environments (security best practices, access reviews).
Experience contributing to a platform/DevEx roadmap that other engineering teams consume as a self-service product.
Our tech stack:
Languages: PHP (Symfony), Node.js (TypeScript), Angular (TypeScript).
Data: PostgreSQL, Redis, MongoDB.
Infra: AWS, Kubernetes, Terraform, Helm, Atlantis
Engineering Tools: Postman, Git, GitHub Copilot, PhpStorm, Grafana, Kibana, Prometheus.
Remote work Tools: Jira, Miro, Google Workspace, Slack.
Development Practices: Pair Programming, Code Reviews, Continuous Integration/Deployment.
What we offer:
A global, inclusive team that’s as supportive as it is ambitious and serious about getting things done
An opportunity to work remotely or in a modern and welcoming office in Riga
Flexible working hours (start your day as late as 11 AM)
Private health insurance
2 extra paid days off to focus on your mental or physical well-being
1 extra paid day off to celebrate a Birthday or any other celebration of your choice
Internal and external learning opportunities
Access to mentorship, internal meetups, and hackathons, both on-site and online
Free and healthy lunch if you work from the Rīga office
Design and order your own merch using our platforms with an employee discount
Exciting team-building events and parties you’ll never forget!
FYUL is the engine that powers on-demand commerce at global scale.
Formed in 2024 through the merger of Printful, Printify, and Snow Commerce, we bring together tech, talent, and infrastructure to help people turn ideas into beautiful products.
From solo creators to entertainment giants, FYUL powers merch that connects with millions, backed by advanced tech, premium production, and global reach.
We're a fast-growing global company working toward powering great brands, great experiences, and great people.
We are an equal-opportunity workplace. We’re committed to diversity and inclusion and make hiring decisions based solely on qualifications, merit, and work experience.
If you think you’d excel in this role, send us your resume in English, showing us why you are the right person for the job.
Interested, but don’t think this is the right fit for you? Feel free to share it with friends and check out other open positions at our career site. We’re always looking for creative and driven minds to join our ever-growing team!
AS Printful Latvia (Reģ. Nr. 40203050078)
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Site Reliability Engineer
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!