Site Reliability / Production Engineer

🏢 Storyteller · all Storyteller jobs
📍 Algeria
💰 EUR 20,000 - 20,000 / annual
📅 Posted 2026-07-23 · via Himalayas
🏷 Site-Reliability-Engineering,Production-Engineering,DevOps,Incident-Response,Operations-Engineering,Senior-Site-Reliability-Engineer,Staff-Site-Reliability-Engineer,Site-Reliability-Engineering-Manager,Site-Reliability-Engineering-Lead
Apply on original site ↗
🇩🇿 Up to EUR 20,000 per year, on a full-time, contractor contract 🌎 Fully remote working from anywhere in Algeria!  🌙 Shared out-of-hours UK coverage, including active evening shifts and weekday overnight pager duty ✨ Exciting high growth product, relied on by leading global brands, particularly within sports 💻 Working with the latest hardware, AI tools, and product workflows. We are looking for hands-on production engineers who can take ownership when live systems need attention: establish the customer impact, investigate the evidence, take safe action and keep the response moving. You will use AI throughout the work, but not as a substitute for judgement. You will be expected to supervise its output, understand the risk of any action and validate that the real customer outcome has recovered. ABOUT US Storyteller is a high-growth B2B SaaS platform that lets companies integrate Stories into their own apps and websites. Popularised by Instagram and Snapchat, Stories help our clients increase engagement, retention and revenue. Our platform includes SDKs for Web, iOS and Android, alongside publishing tools, analytics and advertising support—giving enterprises a complete Stories solution in days. We work with globally recognised sports and media brands, and your work will be used live by millions of people. Our production environment spans Storyteller , Storypilot and the services that support our customers’ live workflows. Reliability is therefore about more than infrastructure: we need to understand when customers are affected, respond quickly, coordinate the right people and improve our systems after every material incident. About the Role We are hiring two Site Reliability / Production Engineers. You will be the first technical response for live incidents during your coverage window. Working closely with Support, you will assess customer impact, investigate the system, take proportionate action and bring in product developers only when their specific knowledge or judgement is genuinely needed. This is not a passive escalation role. You will own the technical response, solve what you reasonably can yourself and make escalations specific and useful. Depending on the incident, you may restart or scale services, roll back deployments, change configuration, repair data, deploy a bounded fix or make a small code change. When incidents are quiet, you will improve the reliability system: reduce alert noise, strengthen customer-outcome monitoring, improve diagnostics, create runbooks and AI Skills, automate repeated work and make our products easier to operate. Working Pattern This role provides out-of-hours production coverage, so the schedule is a core part of the position rather than occasional overtime. - The two hires will share an agreed rota that ensures one engineer is actively working from 17:00-01:00 UK time, seven days a week. - Neither person will work seven days a week; the active shifts will be divided between both hires, with appropriate rest days. - The two engineers will also share weekday pager coverage from 01:00-06:00 UK time. The detailed allocation of active and pager shifts will be explained during the hiring process. - Weekend daytime coverage is provided separately and is not an additional expectation for these roles. - When there are no live incidents, the active shift will be used for reliability-improvement work. - Meetings and collaboration with management and product teams will be arranged within the agreed working pattern. The detailed rota, rest arrangements, leave cover, compensation and on-call terms will be confirmed clearly during the hiring process. Please consider the UK-time evening and overnight requirements carefully before applying. RESPONSIBILITIES Respond to live incidents - Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. - Investigate usi

← All remote jobs