Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)
Job Overview
We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .
The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.
Key Responsibilities & Required Skills
Linux & Bare Metal Infrastructure
- Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
- Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
- Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
- Strong understanding of server hardware, including:
- BIOS/UEFI
- RAID controllers
- Firmware management
- iLO/iDRAC/IPMI
- NICs and SmartNICs
- HBA cards
- Hardware diagnostics and troubleshooting
- Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
AI Factory & GPU Infrastructure
- Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
- Understanding of NVIDIA GPU technologies including:
- A100, H100, H200, B200, or equivalent GPU platforms
- NVIDIA DGX and OEM GPU servers
- GPU provisioning and lifecycle management
- GPU monitoring and performance optimization
- Knowledge of AI Factory architecture and infrastructure requirements
- Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
- Understanding of:
- GPU resource allocation and scheduling
- Multi-GPU systems
- GPU networking requirements
- High-bandwidth, low-latency infrastructure design
- Familiarity with NVIDIA ecosystem technologies such as:
- CUDA
- NCCL
- GPUDirect Storage
- NVIDIA Fabric Manager
- NVIDIA Base Command (preferred)
Enterprise Storage & Data Platforms
- Advanced Linux storage administration:
- LVM
- XFS, EXT4
- NFS
- iSCSI
- Fibre Channel SAN
- Multipath I/O
- Strong hands-on experience with Ceph , including:
- Cluster architecture
- MON, OSD, MDS
- RBD, CephFS, RGW
- Capacity planning
- Performance tuning
- Failure recovery
- Experience with high-performance AI storage platforms such as:
- WEKA
- VAST Data
- Dell PowerScale
- Pure Storage FlashBlade
- NetApp
- Understanding of:
- NVMe-over-Fabrics (NVMe-oF)
- RDMA
- GPUDirect Storage
- Parallel file systems
- AI data pipelines
Networking & Infrastructure
- Strong networking knowledge:
- Bonding
- VLANs
- Routing
- MTU optimization
- DNS
- DHCP
- Experience with high-performance data center networking:
- 100G/200G/400G Ethernet
- RoCE
- RDMA
- Spine-Leaf architectures
- Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
- Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
Operations & Reliability
- Experience with high availability, clustering, and disaster recovery
- Strong troubleshooting skills across:
- Linux operating systems
- Hardware platforms
- GPU infrastructure
- Networking
- Enterprise storage
- Experience supporting mission-critical pro
Get remote developer jobs like this by email
One weekly digest. No spam, unsubscribe anytime.