Infrastructure & Platform Operations Engineer
ABOUT DYSRUPIT
DysrupIT is a consulting-led technology firm. We help mid-market to enterprise businesses solve business problems through technology โ whether that's consulting, execution, managed services, or staff augmentation โ and we take accountability for the outcome. We are dedicated to making a positive impact in the communities we serve.
COMPANY CULTURE
At DysrupIT , success isn't measured in headcount placed or hours billed โ it's measured in outcomes delivered. We're a team that takes ownership of the work, stays curious about problems beyond our immediate scope, and builds relationships meant to grow, not just renew or end. We invest in our people with the training and support they need to grow their careers, and we back a culture where everyone, regardless of role, is encouraged to notice opportunities, ask one more question, and help lead the story for our clients, not just deliver it.
JOB SUMMARY
The Infrastructure & Platform Operations Engineer supports, maintains and improves the infrastructure and production platforms used to deliver Nephos services to customers. The jobholder works across cloud / on-premise infrastructure, Linux, networking, databases, containers, monitoring and enterprise platforms, with a strong initial focus on BAU production support, service stability and knowledge transfer. As capability develops, the role will also contribute to platform change, upgrades, automation and project delivery.
JOB RESPONSIBILITIES
Working hours: 9โ5 UK hours during onboarding and transition, with flexibility for occasional out-of-hours upgrades, changes and major incidents.
- Work as part of Service Operations to support enterprise infrastructure and production platforms, including health monitoring, issue resolution, proactive maintenance, resilience and continuous improvement.
- Administer, support and troubleshoot customer environments, including core infrastructure, storage, networking, IAM awareness, compute resources and managed services relevant to supported platforms.
- Support Linux environments, primarily Ubuntu and Red Hat, using command-line administration and troubleshooting techniques; provide appropriate support for Windows where required.
- Diagnose complex network and connectivity issues across TCP/IP, DNS, routing, firewalls, VPNs, TLS, proxies, load balancers and cloud networking.
- Administer and support PostgreSQL production environments, including configuration, roles and permissions, monitoring, backup and restore, recovery, performance investigation and upgrades.
- Administer and support MongoDB production environments, including health, logs, performance, backup and recovery, upgrades and troubleshooting.
- Support containerised and orchestrated environments using Docker and Kubernetes, including container lifecycle, logs, configuration, volumes, networking, registries, health checks, resource constraints and failure diagnosis.
- Support cloud container services and clusters, including AWS with EKS and ECS where applicable, and contribute to platform configuration, health and capacity management.
- Monitor infrastructure and applications using enterprise monitoring and observability tools, including alert investigation, log analysis, threshold review, root-cause identification and reduction of unnecessary alert noise.
- Use scripting and automation, particularly Python and PowerShell, to improve repeatability, reduce manual effort and support operational tasks; Bash or other transferable scripting experience is also valuable.
- Use Git and GitHub to manage scripts, configuration and operational content, following appropriate version-control practices.
- Support enterprise platform installations, configuration, upgrades, testing, validation and recovery, including data platforms such as BigID and internally developed solutions such as Nephos-developed platforms where customer access is approved.
- Learn and develop operational capability in new or unfamiliar enterp