Site Reliability Engineer

๐Ÿข Second Talent ยท all 19 jobs
๐Ÿ“ Singapore
๐Ÿ“… Posted Aug 23, 2026 ยท via Himalayas
๐Ÿท Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, Systems Operations, Site Reliability Engineer +5 more
Apply on original site โ†—

Cluster Operations & Management

- Manage and maintain container clusters (Kubernetes, Docker) and open-source component clusters (Kafka, Redis, Elasticsearch) across multiple business units

- Ensure optimal performance, scalability, and reliability of distributed systems

Infrastructure Platform Development

- Design, build, and enhance infrastructure operation platforms

- Develop and maintain systems for infrastructure management, CI/CD pipelines, monitoring/alerting, and centralized logging

- Drive platform standardization and automation initiatives

High Availability & Reliability

- Ensure maximum uptime for production services through proactive monitoring and incident response

- Continuously optimize service architecture, deployment strategies, and operational processes

- Implement and maintain SLA/SLO frameworks and reliability engineering practices

Automation & Process Improvement

- Lead the development of automated operations and maintenance systems

- Create self-service tools and workflows to improve team productivity

- Establish best practices for infrastructure such as code and configuration management

Required Qualifications
Experience & Education

- 2+ years of hands-on experience in Systems Operations, DevOps, or Site Reliability Engineering (SRE)

- Bachelor's degree in Computer Science, Engineering, or related technical field preferred

Cloud & Infrastructure

- Experience with public cloud platforms (AWS, Azure, or GCP) is highly valued

- Strong understanding of large-scale internet architecture and distributed systems

- Proven experience with infrastructure monitoring, logging, and observability tools

Technical Skills

- Proficiency in scripting and automation using Shell, Python, or similar languages

- Strong knowledge of containerization technologies (Kubernetes, Docker)

- Hands-on experience operating production-grade container clusters and managing CI/CD pipelines

- Strong familiarity with common infrastructure components: Nginx, MySQL, Redis, Kafka, Elasticsearch

Advanced Networking (Preferred)

- Experience with Service Mesh architectures, Cilium CNI, and eBPF technologies

- Understanding network security, load balancing, and traffic management

- Knowledge of cloud-native networking patterns and best practices

Originally posted on Himalayas

Flights + hotels

This role requires you to be in Singapore. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels โ†’

โ† All remote jobs

Comparing Site Reliability Engineer pay and openings โ€” the live median is $129k?All remote Site Reliability Engineer jobs โ†’Site Reliability Engineer salary data โ†’
Want more like this? Browse every live remote developer role.All remote developer jobs โ†’
Get new developer jobs by email
Daily email, only when there's something new. One click to stop.

Get remote developer jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you