Site Reliability Engineer (Contract, Rotation-Based)

🏢 Invisible Technologies · all Invisible Technologies jobs
📍 India,South Africa
📅 Posted 2026-09-04 · via Himalayas
🏷 Site-Reliability-Engineer,SRE,Infrastructure-Engineering,Production-Operations,DevOps-Engineer,Senior-Site-Reliability-Engineer,Site-Reliability-Operations-Engineer,DevOps-Site-Reliability-Engineer
Apply on original site ↗

About Invisible

Invisible Technologies makes AI work. Our end-to-end AI platform structures messy data, automates digital workflows, deploys agentic solutions, measures outcomes, and integrates human expertise where it matters most.

Our platform cleans, labels, and structures company data so it is ready for AI. It adapts models to each business and adds human expertise when needed, the same approach we have used to improve models for more than 80% of the world’s top AI companies, including Microsoft, AWS, and Cohere.

Our successes span industries, from supply chain automation for Swiss Gear to AI-enabled naval simulations with SAIC, and validating NBA draft picks for the Charlotte Hornets.

Profitable for more than half a decade, Invisible reached $134M in revenue and ranked as the number two fastest growing AI company on the 2024 Inc. 5000. In September 2025, we raised $100M in growth capital to accelerate our mission of making AI actually work in the enterprise and to advance our platform technology.

About The Role

We're looking for Site Reliability Engineers to join a 24/7 incident response rotation supporting a production platform used by one of our key clients. You'll be a first responder for critical incidents — triaging and stabilizing issues at the infrastructure level, working from logs to place a failure to its component, and escalating to the platform team when an issue goes beyond infrastructure-level diagnosis.

We're hiring across a level range. At minimum, you'll own first response: fast, calm triage under time pressure, and good judgment about when to escalate versus resolve. For candidates operating at a more senior level, the role grows to include hardening the systems you're responding to, driving scaling and capacity work, and proposing structural improvements — so the same incidents happen less often over time.
What You’ll Do

- Serve as first responder for production incidents, triaging and stabilizing at the infrastructure level within defined response SLAs (P1/P2 severity)

- Diagnose issues primarily from system logs — Kubernetes, RabbitMQ, and Postgres — to place a failure to the right component before escalating

- Distinguish infrastructure-level failures from application/business-logic failures — for infrastructure-level issues, identify the fix and submit the change yourself; escalate application/business-logic issues to the platform team with clear context

- Participate in an on-call rotation, including off-hours coverage

- Communicate incident status clearly to stakeholders during active incidents, and hand off cleanly to the team that owns resolution

What We Need

- Solid hands-on experience with Kubernetes, RabbitMQ, and PostgreSQL in an enterprise setting, ideally within Financial Services

- Strong working knowledge of Azure; familiarity with GCP or AWS is a plus

- Comfort diagnosing an unfamiliar system primarily from its logs rather than its source — this matters more than deep expertise in any one tool

- Experience troubleshooting production systems under time pressure, with sound judgment about severity and escalation

- Clear, calm communication during live incidents

Engagement Type & Schedule

This is a contractor role structured around a coverage rotation rather than a standard full-time schedule. The commitment breaks down as:

- 10+ hours per week during regular business hours

- Rotational weekend coverage: 8 hours on Saturday and 8 hours on Sunday, every other weekend

- ~72+ hours per month total, combining weekday and rotational weekend coverage

*Please note: this role is an hourly, contract-based position and is not eligible for bonus or equity compensation*

You can find more information about our geographic pay tiers here. During the interview process, your Invisible Talent Acquisition Partner will confirm which tier applies to your location. For candidates outside the U.S., compensation is adjusted to reflect local market conditions and cost of living.

← All remote jobs

Similar for you