Data Site Reliability Engineer (SRE)

🏢 General Dynamics Information Technology · all General Dynamics Information Technology jobs
📍 United States
💰 USD 111,155 - 150,385 / annual
📅 Posted 2026-08-30 · via Himalayas
🏷 Site-Reliability-Engineering,Data-Engineering,Cloud-Engineer,DevOps-Engineer,IT-Infrastructure-And-Operations,Data-Reliability-Engineer-Jobs,Database-Site-Reliability-Engineer,Site-Reliability-Engineer,Site-Reliability-Engineering-(SRE),Site-Reliability-Operations-Engineer,Database-Reliability-Engineer,Data-Reliability-Engineering
Apply on original site ↗

Type of Requisition:
Regular
Clearance Level Must Currently Possess:
None
Clearance Level Must Be Able to Obtain:
None Public Trust/Other Required:
BI Full 6C (T4)
Job Family:
IT Infrastructure and Operations Job Qualifications:
Skills:
CI/CD, Containerization, Structured Query Language (SQL) Development Certifications:
None Experience:
5 + years of related experience US Citizenship Required:
No
Job Description:

Seize your opportunity to make a personal impact supporting the Case Management Modernization (CMM) Program. The CMM program is an initiative to support the Administrative Office of the US Courts (AO) in developing a modern cloud-based solution to support all 204+ federal courts across the United States.

GDIT is your place to make meaningful contributions to challenging projects and grow a rewarding career. The Data Site Reliability Engineer (SRE) will work as part of the CMM Data Modernization and Governance team responsible for delivering integrated data governance, engineering, data platform, reporting, analytics, and Artificial Intelligence (AI)/Machine Learning (ML) capabilities that support operational decision-making and fulfill the AO's data and analytics objectives in support of the CMM program.

The successful candidate will be responsible for providing technical leadership for the day-to-day operational support, reliability, performance, and continuous improvement of the CMM data platforms, pipelines, applications, and analytics services. This role ensures that data services remain secure, available, reliable, and aligned with established service levels, data governance standards, architecture principles, and operational procedures.

THE DATA SITE RELIABILITY ENGINEER (SRE) WILL EXECUTE THE FOLLOWING RESPONSIBILITIES

-
Provide comprehensive real-time monitoring, incident and event management, capacity planning, and operational reporting to support application deployments, maintain system health, predict demand, and align cloud operations with evolving business and security objectives.

-
Maintain and audit user roles and responsibilities in cloud environments.

-
Integrate Single Sign On (SSO), Multi-Factor Authentication (MFA) and group identity management managed through the Judiciary Enterprise Network Information Exchange (JENIE) for enforcing least privilege access.

-
Adhere to guidelines prescribed by the Government and continuously assess and improve credential management processes for all user credentials.

-
Provide Disaster Recovery (DR) and Continuity of Operations (COOP) options. This must include high-availability options, including fault-tolerant and automated failover designs.

-
Integrate DevSecOps tools and processes seamlessly with enterprise systems (Integrated Development Environments (IDEs), ticketing, monitoring, etc.) to avoid fragmentation and ensure unified security posture.

-
Provide and manage a centralized secrets management system with automated rotation, access logging, and policy enforcement to securely store, manage, and control access to sensitive information and to prevent unauthorized access and data breaches for any administrative user account.

-
Integrate security tools (example: SAST, DAST, SCA, CSPM) into pipelines for continuous assessment and remediation.

-
Implement unified, automated, continuous monitoring (24/7/365) systems and tools for security, performance, and compliance across all environments, leveraging dashboards and alerting for real-time visibility. Provide supplemental monitoring of event response activities beyond normal business hours (7a.m – 6p.m Eastern Time). Systems and tools shall capture data without including a required response to alerts.

-
Ensure automated generation and management of Software Bill of Materials (SBOM) for all deployed artifacts, supporting transparency and compliance.

-
Provide diagnostics, metrics’ gathering, and performance tuning services.

-
Provide canary release function for end-user testi

← All remote jobs

Similar for you