Senior DevOps / SRE Engineer

๐Ÿข MLabs ยท all MLabs jobs
๐Ÿ“ United States
๐Ÿ’ฐ USD 120,000 - 150,000 / annual
๐Ÿ“… Posted 2026-07-05 ยท via Himalayas
๐Ÿท Backend-Engineering,DevOps,Site-Reliability-Engineering,Platform-Engineering,Infrastructure-Engineering,Senior-Staff-DevOps-Engineer,Sr-Staff-DevOps-Engineer,Senior-DevOps,Senior-DevOps-Platform-Engineer,DevOps-Engineer,Cloud-DevOps-Engineer
Apply on original site โ†—

Staff Software Engineer - Backend & AI Infra

Remote | Full-time

Location: Based in US to GMT timezones

Compensation: Competitive Compensation Package

Our client is a high-growth technology firm. They are seeking a Staff Software Engineer to spearhead two critical domains: the core agent runtime and backend infrastructure powering a high-frequency trading fleet, and the comprehensive migration of model hosting and agent deployment to in-house, proprietary infrastructure.

This is a foundational, high-impact building role. The successful candidate will design and implement the backend services, runtime engines, and deployment systems that enable a fleet of autonomous agents to operate with superior speed, reliability, and intelligence. By moving away from third-party LLM providers and hosted platforms, this role will establish the sovereign infrastructure necessary for the next generation of autonomous financial software.
Key Responsibilities
Agent Runtime & Backend Development

-
Plugin Runtime Ownership: Lead the evolution of the per-agent process, migrating from a distributed Go/Python hybrid to a centralized, high-performance Go service utilizing Postgres state and real-time websocket price feeds.

-
Rules Engine Engineering: Build a YAML-configurable "Scanner Gateway" to bridge signal production and execution, allowing for complex scoring and filtering without direct code manipulation.

-
Advanced Execution Systems: Develop and maintain the RatchetStop Backend , a centralized profit-trailing service capable of sub-second evaluation and websocket-based order execution to protect capital even when agents are offline.

-
Data & Connectivity: Manage the Model Context Protocol (MCP) server bridging agents to platform tools, and oversee a high-throughput data pipeline (Redis, Postgres, ClickHouse) for real-time market intelligence ingestion.

Model & Agent Hosting Migration

-
Infrastructure Sovereignty: Lead the technical execution of migrating agents from third-party platforms to a custom-built, Senpi-hosted environment featuring isolated workspaces and state persistence.

-
Model Serving: Evaluate and implement the transition from external LLM APIs (Anthropic, Google) to self-hosted inference, optimizing for telemetry capture and performance.

-
Telemetry & Feedback Loops: Architect systems to capture every agent decision and score, creating a self-reinforcing loop where the fleet learns and improves from collective performance data.

-
Deployment Pipelines: Build robust CI/CD pipelines for zero-downtime rollouts, ensuring that updates to scanner logic or runtime patches do not interrupt active market positions.

Infrastructure & Operations

-
System Reliability: Design monitoring and alerting frameworks to detect agent failures, state corruption, or authentication expirations before they impact financial performance.

-
Cloud Orchestration: Manage AWS/EKS environments using Infrastructure-as-Code (IaC).

-
Incident Response: Own the operational health of the fleet, acting as the primary responder for high-stakes trading system incidents.

I Interview Process

-
Founder / CEO Interview: Introduction to the vision and strategic goals.

-
Take-Home Test: A practical assessment of technical design and coding capabilities.

-
Technical Interview: A deep dive into systems architecture and engineering expertise.

-
Final Interview: Cultural alignment and final technical synthesis.

Requirements

- Technical Essentials

-
Expert Backend Engineering: Proficiency in writing production-grade code in Go , Python , and Node.js/TypeScript (Go is strongly preferred for runtime services).

-
Startup Experience: A proven track record of building complex backend services (APIs, job scheduling, distributed systems) from scratch in a fast-paced environment.

-
Real-Time Systems: Deep understanding of low-latency environments, websocket management, and sub-second condition evaluation.

-
Database

โ† All remote jobs