Azure CloudOps Engineer

๐Ÿข Embrace Software Inc ยท all Embrace Software Inc jobs
๐Ÿ“ United States
๐Ÿ“… Posted 2026-07-29 ยท via Himalayas
๐Ÿท CloudOps-Engineer,Azure-Cloud-Engineer,DevOps-Engineer,Platform-Engineer,Site-Reliability-Engineer,Cloud-Operations-Engineer,Azure-Engineer,Azure-Platform-Engineer,Azure-Infrastructure-Engineer
Apply on original site โ†—

This is a remote position.
We are looking for a CloudOps Engineer to operate and continuously improve the reliability, security, scalability, observability, and cost efficiency of our Azure-hosted SaaS products. Our products are deployed across development, QA, staging, and production environments, with infrastructure managed through Terraform and CI/CD automated through GitHub Actions. โ€‹

This role will work closely with engineering teams to ensure our SaaS platforms and AI-enabled solutions are deployed consistently, monitored effectively, secured properly, and operated reliably in production.

Environment and Technology Context
- Microsoft Azure-hosted SaaS products across dev, QA, staging, and production environments.

- Terraform for infrastructure as code and repeatable environment provisioning.

- GitHub Actions for application and infrastructure CI/CD workflows.

- Azure services including Static Web Apps, Container Apps, PostgreSQL, Storage Accounts, SignalR, Service Bus, Azure AI Foundry, Speech-to-Text services, Azure Arc, and related services.

- AI-enabled product capabilities including STT workloads, LLM integrations, AI service endpoints, quotas, usage monitoring, latency monitoring, and cost controls.

Key Responsibilities

Cloud Infrastructure Operations
- Manage and support Azure cloud infrastructure across dev, QA, staging, and production environments.

- Maintain operational health of Azure services including Static Web Apps, Container Apps, PostgreSQL, Storage Accounts, SignalR, Service Bus, Azure AI Foundry, Azure Arc, and related platform services.

- Ensure cloud resources are provisioned, configured, monitored, maintained, and retired according to company standards.

- Support environment setup for new products, customers, integrations, and internal initiatives.

- Identify and resolve infrastructure issues affecting performance, reliability, availability, or security.

Terraform and Infrastructure as Code
- Build, maintain, and improve Terraform modules and environment configurations.

- Ensure infrastructure changes are version-controlled, peer-reviewed, tested, approved, and repeatable.

- Manage Terraform state, workspaces, variables, secrets integration, and deployment workflows.

- Detect and resolve configuration drift between Terraform and deployed Azure resources.

- Standardize naming conventions, tagging, resource group structure, environment isolation, and module patterns.

- Support scalable provisioning of new SaaS environments using reusable infrastructure templates.

GitHub Actions and CI/CD
- Build, maintain, and troubleshoot GitHub Actions workflows for application and infrastructure deployments.

- Support CI/CD pipelines for multiple SaaS products and environments.

- Implement deployment promotion flows from development to QA to staging to production.

- Add deployment safeguards such as environment protection rules, approvals, rollback procedures, validation checks, release gates, and audit trails.

- Manage pipeline secrets, service principals, managed identities, and secure deployment credentials.

- Improve build and deployment reliability, speed, traceability, and auditability.

AI Service Operations
- Operate and monitor Azure AI services, including Azure AI Foundry and Speech-to-Text workloads.

- Support production operations for LLM-based integrations and AI-enabled product features.

- Monitor AI service availability, latency, quota usage, token consumption, API failures, throttling, and cost.

- Help define operational standards for AI workloads, including access control, logging, alerting, failover, usage governance, and provider disruption handling.

- Work with engineering teams to troubleshoot AI service issues, integration failures, degraded model responses, or provider-side service disruptions.

- Support secure handling of AI-related secrets, endpoints, keys, managed identities, and private network access where applicable.

Monitoring, Alerting, and

โ† All remote jobs