Staff Software Engineer, Alerting Platform
About Command|Link
Command|Link is a global SaaS Platform providing network, voice services, and IT security solutions, helping corporations consolidate their core infrastructure into a single vendor and layering on a proprietary single pane of glass platform. Command|Link has revolutionized the IT industry by tackling the problems our competitors create. In recognition for our unprecedented innovation and dedication, Command|Link was recognized as the SD-WAN Product of the Year, ITSM Visionary Spotlight, UCaaS Product of the Year, NaaS Product of the Year, Supplier of the Year, and the AT&T Strategic Growth Partner. Command|Link has built the only IT platform for scale that solves ISP vendor sprawl and IT headaches. We make it easy for our customers to get more done, maximize uptime and improve the bottom line.
Learn more about us here!
This is a 100% remote position
About your new role:
Command|Alert is CommandLink 's signal-processing core: the engine that turns raw security, monitoring, and customer-defined telemetry into alerts customers actually trust. Alert fatigue and noise are the top complaint across every competitor in this space, and this role exists to make sure our alerts are the ones people don't tune out.
As a Staff Software Engineer on Command|Alert, you'll operate across the two to three teams that touch alerting, from rule evaluation and anomaly detection through delivery and downstream notification. You'll drive the org's most consequential decisions on how the correlation and alerting engine is architected, dive into whichever team or project needs your depth, and balance near-term reliability work with the long-term technical foundation the product is built on.
This is a role for someone who reasons fluently across data: taking in security tooling, monitoring telemetry, syslog, OpenTelemetry, and L2-L4 network protocols, deriving real network and system topologies and the dependencies between them, and using that context to make LLM-driven reasoning over correctness, troubleshooting, and remediation actually work.
Key Responsibilities:
- Set the architecture for how Command|Alert evaluates rule-based thresholds, ML anomaly scores, and correlation logic to produce high-fidelity alerts, across both global alerts we define and alerts customers define themselves.
- Own the reliability of the alerting pipeline end to end: from OpenSearch alert evaluation through Kafka delivery via OpenSearch callbacks to downstream notification, including idempotency guarantees and soak-tested behavior under sustained load.
- Drive the correlation strategy that turns diverse sources (security tooling, monitoring telemetry, syslog, OpenTelemetry, L2-L4 network protocols) into usable network and system topologies and their dependencies, and make that context usable for LLM reasoning over investigations and remediation.
- Lay the technical groundwork for generating alert definitions from the normalized data model using LLMs.
- Jump into any team or workstream across Command|Alert that needs architectural guidance, unblocking others and raising the bar on how the system is built.
- Balance strategic bets (new correlation and detection capability) against the long-term foundation the alerting engine needs to hold up at scale.
- Mentor engineers across the teams you touch, and represent Command|Alert's technical direction to stakeholders outside engineering.
- Takes on additional responsibilities and projects as needed to support the success of the team and organization.
What you'll need for success:
Required
- Demonstrated experience designing, building, or operating high-reliability alerting or notification systems in production, including running rule-based and ML-based detection logic against real traffic at scale.
- Strong Kafka experience, particularly producing and consuming event streams for downstream delivery.
- A working command of telemetry and protocol data: security tooling output, monito