Sr. Platform Engineer

🏢 Alkami Technology, Inc. · all Alkami Technology, Inc. jobs
📍 United States
💰 USD 145,000 - 165,000 / annual
📅 Posted 2026-08-09 · via Himalayas
🏷 Platform-Engineering,Site-Reliability-Engineering,DevOps-Engineer,Software-Engineering,Cloud-Infrastructure-Engineering,Senior-Platform-Engineer,Senior-Platform-Engineering,Senior-Principal-Platform-Engineer,Senior-Developer-Platform-Engineer,Senior-DevOps-Platform-Engineer
Apply on original site ↗
Alkami is the digital sales and service platform provider for U.S. banks and credit unions. Our unified Platform integrates onboarding, digital banking, and data and marketing—each solution can stand alone, but together they deliver more—to help institutions onboard, engage, and grow relationships. As the future shifts toward Anticipatory Banking, we help data-informed bankers meet the moment with technology that drives action. Founded in 2009, we continue to be recognized for our intentional culture and tremendous growth (Best Place to Work in Fintech; Best & Brightest to Work For Nationally; and Comparably’s Best Company Culture, Best Career Growth, Best Engineering Team, and Best Places to Work in Dallas, among others). We’re building a culture where each Alkamist can perform to their highest potential, and we’re always on the lookout for the best and brightest minds. If you’re ready to experience the power of alchemy - transforming the ordinary into the extraordinary - come join one of the fastest growing SaaS companies in the U.S. As a remote-first company, most of our positions can be remote in the US, except for key roles, which will be indicated in the Job Title. Follow us on Glassdoor and LinkedIn! The Sr Platform Engineer is responsible for locating, making visible, and remediating sources of unreliability in the MANTL platform, including correctness problems that surface under failure conditions. This role works directly in the platform's application codebase (TypeScript) and its container-native deployment environment, combining application-engineering skill with reliability-engineering practice. Much of the work is reactive, investigating and resolving issues as they surface, balanced against a standing roadmap of known reliability risks the team has identified and prioritized ahead of time, independent of feature-delivery timelines. This role partners with, but is organizationally and functionally distinct from, both Cloud Infrastructure Engineering and product application engineering teams, focusing specifically on reliability concerns that span or fall between those domains. Essential Duties & Responsibilities - Investigate, troubleshoot, and resolve reliability issues within MANTL platform application code, including microservice communication failures and correctness issues that emerge under failure conditions - Identify and address failure modes across the platform's third-party and internal system integrations, developing resilience strategies suited to each integration's specific behavior - Design, configure, and maintain monitoring, dashboards, and alerting (Datadog preferred) to increase visibility into platform health and surface emerging issues before they become incidents - Implement and extend distributed tracing across microservices to accelerate root-cause identification for cross-service failures - Diagnose and remediate application performance issues, including caching strategy, inefficient code paths, and query performance - Harden platform and application components against known failure modes through fault-injection and resilience testing, implementing defensive design patterns to prevent recurrence - Build, maintain, and troubleshoot CI/CD build pipelines (GitHub Actions) supporting deployment of the platform - Deploy and troubleshoot container-native (Kubernetes) workloads as part of diagnosing and resolving platform reliability issues - Maintain and execute against a roadmap of known reliability risks, independent of feature-delivery timelines - Create and maintain documentation and runbooks covering platform reliability issues, root causes, and remediations - Act as an escalation point for complex platform reliability issues, partnering with Cloud Infrastructure Engineering and application engineering teams on issues that cross domain boundaries - Contribute to defining reliability targets (SLOs/SLIs) for key platform services Recommended Experience & Educatio

← All remote jobs