Please mention DailyRemote when applying
🇹🇳 Up to EUR 25,000 per year, on a full-time, contractor contract
🌎 Fully remote working from anywhere in Tunisia!
🌙 Shared out-of-hours UK coverage, including active evening shifts and weekday overnight pager duty
✨ Exciting high growth product, relied on by leading global brands, particularly within sports
💻 Working with the latest hardware, AI tools, and product workflows.
We are looking for hands-on production engineers who can take ownership when live systems need attention: establish the customer impact, investigate the evidence, take safe action and keep the response moving.
You will use AI throughout the work, but not as a substitute for judgement. You will be expected to supervise its output, understand the risk of any action and validate that the real customer outcome has recovered.
ABOUT US
Storyteller is a high-growth B2B SaaS platform that lets companies integrate Stories into their own apps and websites. Popularised by Instagram and Snapchat, Stories help our clients increase engagement, retention and revenue.
Our platform includes SDKs for Web, iOS and Android, alongside publishing tools, analytics and advertising support—giving enterprises a complete Stories solution in days. We work with globally recognised sports and media brands, and your work will be used live by millions of people.
Our production environment spans Storyteller, Storypilot and the services that support our customers’ live workflows. Reliability is therefore about more than infrastructure: we need to understand when customers are affected, respond quickly, coordinate the right people and improve our systems after every material incident.
About the Role
We are hiring two Site Reliability / Production Engineers. You will be the first technical response for live incidents during your coverage window.
Working closely with Support, you will assess customer impact, investigate the system, take proportionate action and bring in product developers only when their specific knowledge or judgement is genuinely needed.
This is not a passive escalation role. You will own the technical response, solve what you reasonably can yourself and make escalations specific and useful. Depending on the incident, you may restart or scale services, roll back deployments, change configuration, repair data, deploy a bounded fix or make a small code change.
When incidents are quiet, you will improve the reliability system: reduce alert noise, strengthen customer-outcome monitoring, improve diagnostics, create runbooks and AI Skills, automate repeated work and make our products easier to operate.
Working Pattern
This role provides out-of-hours production coverage, so the schedule is a core part of the position rather than occasional overtime.
The detailed rota, rest arrangements, leave cover, compensation and on-call terms will be confirmed clearly during the hiring process. Please consider the UK-time evening and overnight requirements carefully before applying.
RESPONSIBILITIES
Respond to live incidents
Coordinate the right response
Improve the reliability system
QUALIFICATIONS
What we're looking for
Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response. It is not an automatic requirement: we will also consider candidates who demonstrate exceptional ownership, judgement, technical aptitude, learning velocity and performance in the practical assessment.
You do not need experience with every technology in our stack, a previous SRE job title, people-management experience or the ability to recall every command without AI assistance. The ability to learn an unfamiliar environment, act safely and validate your work matters more than matching a long technology checklist.
Nice to have
RECRUITMENT PROCESS
We keep the process straightforward, practical and respectful of your time.
1. Hiring Manager Conversation (20-30 mins)
A short call to get to know you, talk through the working pattern and answer your questions.
2. Paid Take-home Task (~60-90 mins)
A small, bounded production-incident exercise using evidence such as a Support report, alerts, logs, metrics, deployment history, code and an imperfect runbook. We compensate you for completing it regardless of the outcome.
You are encouraged to use AI. We are interested in how you establish impact, investigate and revise hypotheses, choose a safe response, validate the outcome and communicate the incident - not in your ability to reproduce commands or syntax from memory.
3. Task Review and CTO Interview (60-75 mins)
We will review your submission together, explore the decisions and trade-offs you made, and discuss how you supervised AI-generated analysis or changes. You will also meet Dave, our CTO, and talk about production judgement, escalation, validation, communication and how you improve the system after an incident.
And that's it.
----------------------
Privacy Notice
We process your personal data for recruitment purposes in line with UK data protection law. AI tools may assist in reviewing applications, but decisions are made by our team. We retain data only as necessary for recruitment and compliance. You can request access or deletion of your data at any time by emailing careers@getstoryteller.com.
Stop the endless job search. Our AI finds and applies to the best jobs for you.
Discover remote opportunities in Software Development
Answer easy questions
200,000+ jobs across 15+ categories
Get your best job matches
Only hand-screened, legit jobs
Find a remote job faster
No ads, scams, or junk
“ I was the first applicant for a remote marketing position that got listed on the company website the same day I applied. Had an interview within 48 hours!