Expert Automation & Observability Engineer

🏢 Ensono · all 32 jobs
📍 United States
💰 USD 140,000 - 180,000 / annual
📅 Posted Sep 20, 2026 · via Himalayas
🏷 Observability Engineering, Site Reliability Engineering, Automation Engineering, Cloud Engineer, Monitoring Operations, DevOps Automation Engineer +8 more
Apply on original site ↗

At Ensono , our Purpose is to be a relentless ally, disrupting the status quo and unleashing our clients to Do Great Things ! We enable our clients to achieve key business outcomes that reshape how our world runs. As an expert technology adviser and managed service provider with cross-platform certifications, Ensono empowers our clients to keep up with continuous change and embrace innovation.

We can Do Great Things because we have great Associates. The Ensono Core Values unify our diverse talents and are woven into how we do business. These five traits are the key to achieving our purpose:

Honesty, Reliability, Curiosity, Collaboration, and Passion.
About the role and what you'll be doing:

We are seeking an Expert Observability Engineer to serve as the strategic technical lead and architect for our enterprise Observability, APM, and Telemetry ecosystems. You will lead the transformation from decentralized, reactive monitoring to a unified, automated, and proactive observability framework. Operating across hybrid cloud, Kubernetes, and legacy environments, you will design scalable architectures, drive Site Reliability Engineering (SRE) practices, and lead First-of-a-Kind (FOAK) technology implementations to ensure maximum service reliability.

We want all new Associates to succeed in their roles at Ensono . That's why we've outlined the job requirements below. To be considered for this role, it's important that you meet all Required Qualifications. If you do not meet all of the Preferred Qualifications, we still encourage you to apply.
Core Responsibilities
1. Enterprise Architecture & Strategy

- Architect and govern a unified observability framework covering metrics, logs, traces, and events using IBM Instana, Grafana, OpenTelemetry, Telegraf, and InfluxDB .

- Lead First-of-a-Kind (FOAK) implementations—evaluating new observability tech and converting them into secure, repeatable, production-ready patterns.

- Define enterprise standards for telemetry pipelines, data retention, high-cardinality controls, and observability cost management.

2. SRE & Service Reliability

- Define and govern Service Level Indicators (SLIs), Objectives (SLOs), and error budgets.

- Serve as the senior technical escalation point, leading major P1/P2 incident war rooms and conducting evidence-based Root Cause Analysis (RCA).

- Drastically reduce MTTD/MTTR and alert noise through event correlation, dynamic thresholds, and dependency mapping.

3. Platform Engineering & Automation

- Drive Observability-as-Code and infrastructure automation using Ansible, Terraform, Python, and GitOps .

- Automate the deployment, configuration, and self-healing workflows for monitoring agents and telemetry collectors.

- Integrate observability platforms seamlessly with ITSM (ServiceNow), Netcool, and CI/CD pipelines.

4. Cloud-Native & Kubernetes Observability

- Design deep observability for Docker, Kubernetes, microservices, and multi-cloud environments (Azure/AWS/GCP).

- Correlate application APM telemetry with Kubernetes control planes, pods, nodes, and infrastructure dependencies.

- Ensure secure-by-design telemetry pipelines (RBAC, TLS, secrets management, and image scanning).

5. Technical Leadership & Transition Management

- Lead complex Knowledge Transfer (KT) programs, vendor transitions, and operational readiness handovers for global 24x7 teams.

- Mentor cross-functional engineering teams and influence enterprise technology roadmaps.

Required Technical Stack

-
Observability & APM: IBM Instana, Grafana (Enterprise & Alloy), Prometheus, OpenTelemetry, Telegraf, InfluxDB.

-
Legacy/Traditional Monitoring: SolarWinds, Netcool, Elastic/Splunk.

-
Cloud & Containerization: Kubernetes, Docker, OpenShift, AWS/Azure/GCP.

-
Infrastructure: Linux (RHEL), Windows Server, VMware, Citrix VDI, load balancers, and edge proxies.

-
Automation & DevOps: Ansible, Terraform, Python, Bash, Webhooks, CI/CD (GitHub Actions/GitLab/Jenkins).

-

Flights + hotels

This role requires you to be in the United States. If that means relocating or flying in, it is worth checking fares before you commit to a start date.

Compare flights and hotels →

← All remote jobs

Want more like this? Browse every live remote developer role.All remote developer jobs →
Get new developer jobs by email
Daily email, only when there's something new. One click to stop.

Get remote developer jobs like this by email

10 hand-picked jobs, one email a day. No spam, unsubscribe anytime.

Similar for you