Site Reliability Engineer
About IXOPAY
IXOPAY is the enterprise-grade global payment infrastructure platform built for the era of agentic commerce. We equip merchants and enterprises with AI-driven intelligence, payment orchestration, advanced tokenization, and the tools to optimize every stage of the payments journey.
From intelligent routing and compliance to customizable modules and enterprise-scale orchestration, IXOPAY helps businesses integrate faster, improve payment performance, and expand globally with confidence.
At IXOPAY , our people are our greatest strength. Our collaborative culture is built on shared values that guide how we work, innovate, and support our customers.
About the Role
We're looking for a Site Reliability Engineer to help build and operate the resilient, scalable infrastructure that powers IXOPAY 's global payment platform.
In this hands-on engineering role, you'll combine software engineering and platform expertise to improve the reliability, performance, and scalability of our cloud-native systems. Working closely with Engineering, Security, Product, and Infrastructure teams, you'll automate operational processes, strengthen platform resilience, and ensure our services meet the high availability and performance expectations of enterprise customers.
This is an opportunity to help shape the reliability engineering practices that support one of the world's leading intelligent payments platforms.
What You'll Do
Build Reliable Platforms
-
Design, build, and maintain software and automation that improves the reliability, scalability, and operational efficiency of IXOPAY 's platform.
-
Develop infrastructure and platform services using cloud-native technologies across Microsoft Azure and Amazon Web Services (AWS).
-
Contribute to system architecture, platform management, capacity planning, testing, and release processes that enable secure, resilient growth.
-
Balance delivery speed with reliability by supporting well-defined service level objectives (SLOs) and operational best practices.
Improve Operational Excellence
-
Proactively monitor the health, availability, latency, performance, and capacity of customer-facing applications and infrastructure.
-
Lead troubleshooting and recovery efforts for production incidents, driving resolution through structured problem solving and strong ownership.
-
Conduct root cause analyses for critical incidents and implement long-term improvements that reduce operational risk.
-
Create and maintain technical documentation, operational runbooks, and standardized procedures that improve team effectiveness.
Automate and Optimize
-
Identify opportunities to automate routine operational tasks, infrastructure management, and incident response workflows.
-
Build sustainable systems through infrastructure automation, configuration management, and continuous improvement initiatives.
-
Gather and analyze platform and application metrics to improve performance, reliability, and fault detection.
-
Support disaster recovery planning, resilience testing, and ongoing platform optimization.
Collaborate Across Teams
-
Partner with Software Engineering teams throughout the development lifecycle to improve reliability, observability, and operational readiness.
-
Participate in Agile planning and delivery, contributing to projects that enhance platform capabilities and operational excellence.
-
Share technical knowledge, mentor teammates, and continuously expand your own expertise in cloud infrastructure and reliability engineering.
What You'll Bring
Required
-
Experience in Site Reliability Engineering, Platform Engineering, DevOps, or a related infrastructure engineering role.
-
Experience supporting cloud-native applications in Microsoft Azure, Amazon Web Services (AWS), or both.
-
Strong understanding of Linux administration, system troubleshooting, and performance tuning.
-
Experience with infrastructure as code and configuration management tools such