Site Reliability Engineer 3

🏢 Granicus · all Granicus jobs
📍 India
📅 Posted 2026-08-03 · via Himalayas
🏷 Site-Reliability-Engineer,AIOps-Engineer,DevOps-Engineer,Cloud-Reliability-Engineer,Production-Engineer,Senior-Site-Reliability-Engineer,DevOps-Site-Reliability-Engineer,Staff-Site-Reliability-Engineer-(SRE),Principal-Site-Reliability-Engineer,Site-Reliability-Engineering-Lead
Apply on original site ↗

The Company
Serving the People Who Serve the People

Granicus is driven by the excitement of building, implementing, and maintaining technology that is transforming the Govtech industry by bringing governments and its constituents together. We are on a mission to support our customers with meeting the needs of their communities and implementing our technology in ways that are equitable and inclusive. Granicus has consistently appeared on the GovTech 100 list over the past 5 years and has been recognized as the best companies to work on BuiltIn.

Over the last 25 years, we have served 5,500 federal, state, and local government agencies and more than 300 million citizen subscribers power an unmatched Subscriber Network that use our digital solutions to make the world a better place. With comprehensive cloud-based solutions for communications, government website design, meeting and agenda management software, records management, and digital services, Granicus empowers stronger relationships between government and residents across the U.S., U.K., Australia, New Zealand, and Canada. By simplifying interactions with residents, while disseminating critical information, Granicus brings governments closer to the people they serve—driving meaningful change for communities around the globe.

Want to know more? See more of what we do here.
Job Summary
Job Description:

Granicus is seeking a Site Reliability Engineer 3 (SRE) with strong AIOps, automation, and AI proficiency to modernize reliability engineering through observability, intelligent incident response, and responsible AI-assisted operations. In this role, you will improve service reliability, reduce operational toil, accelerate incident response, and help build scalable, resilient platforms supporting traditional, cloud-native, and AI/ML-powered workloads. The role will also help operationalize AI-enabled SRE practices such as alert intelligence, assisted root-cause analysis, runbook automation, telemetry summarization, and governed self-healing workflows with appropriate human approval and audit controls. You will be expected to implement practical AIOps capabilities across observability, incident response, automation, and operational knowledge workflows, turning AI/ML insights into production-ready reliability improvements.
What Your Impact Will Look Like

- End-to-end reliability for production systems: on-call, incident response, postmortems, SLO/SLI ownership

- Building and maintaining observability pipelines (metrics, logs, traces) with AI-assisted anomaly detection layered on top and implementing AIOps pipelines for event ingestion, enrichment, deduplication, correlation, and noise reduction

- Designing automation that reduces toil — with a clear bias toward AI-augmented runbooks over static scripts and implementing AIOps-driven remediation workflows with approvals, rollback logic, and audit trails

- Driving alert-noise reduction using correlation/ML techniques, not just threshold tuning

- Partnering with engineering teams to embed reliability and AI-assisted diagnostics into the SDLC

- Demonstrated production use of LLMs/AI agents for SRE workflows — e.g., automated log triage, RCA drafting, runbook generation, or incident summarization — not just "I used Copilot to write YAML"

- Experience building or integrating AIOps capabilities: anomaly detection, alert correlation/clustering, predictive capacity signals including hands-on implementation using observability platforms, ML-based signal processing, incident enrichment, and ChatOps/ticketing integrations

- Working knowledge of prompt engineering for operational use cases (structured outputs, tool-use/function calling, retrieval-augmented context from runbooks/CMDB)

- Comfort evaluating AI output critically — can articulate where an LLM's suggested fix or RCA was wrong and why, not just accept it

- Familiarity with agentic frameworks or MCP-style tool integration (connecting LLMs to ticketing, observabilit

← All remote jobs