Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

🏒 uvation · all uvation jobs (2)
πŸ“ India
πŸ“… Posted 2026-09-08 Β· via Himalayas
🏷 Linux-Infrastructure-Engineering,Bare-Metal-Infrastructure,AI-Factory-Infrastructure,GPU-Infrastructure-Engineering,Storage-Engineering,Linux-Infrastructure-Engineer,Linux-Infrastructure-Specialist,Infrastructure-Linux-Unix-Analyst,Senior-AI-Infrastructure-Engineer,AI-Infrastructure-Engineer,AI-ML-Infrastructure-Engineer,Machine-Learning-Infrastructure-Engineer,Infrastructure-Platform-Engineer,Infrastructure-Engineer
Apply on original site β†—

Job Overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills

Linux & Bare Metal Infrastructure

- Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)

- Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management

- Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments

- Strong understanding of server hardware, including:
- BIOS/UEFI

- RAID controllers

- Firmware management

- iLO/iDRAC/IPMI

- NICs and SmartNICs

- HBA cards

- Hardware diagnostics and troubleshooting

- Experience designing, implementing, and supporting enterprise Linux infrastructure at scale

AI Factory & GPU Infrastructure

- Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads

- Understanding of NVIDIA GPU technologies including:
- A100, H100, H200, B200, or equivalent GPU platforms

- NVIDIA DGX and OEM GPU servers

- GPU provisioning and lifecycle management

- GPU monitoring and performance optimization

- Knowledge of AI Factory architecture and infrastructure requirements

- Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads

- Understanding of:
- GPU resource allocation and scheduling

- Multi-GPU systems

- GPU networking requirements

- High-bandwidth, low-latency infrastructure design

- Familiarity with NVIDIA ecosystem technologies such as:
- CUDA

- NCCL

- GPUDirect Storage

- NVIDIA Fabric Manager

- NVIDIA Base Command (preferred)

Enterprise Storage & Data Platforms

- Advanced Linux storage administration:
- LVM

- XFS, EXT4

- NFS

- iSCSI

- Fibre Channel SAN

- Multipath I/O

- Strong hands-on experience with Ceph , including:
- Cluster architecture

- MON, OSD, MDS

- RBD, CephFS, RGW

- Capacity planning

- Performance tuning

- Failure recovery

- Experience with high-performance AI storage platforms such as:
- WEKA

- VAST Data

- Dell PowerScale

- Pure Storage FlashBlade

- NetApp

- Understanding of:
- NVMe-over-Fabrics (NVMe-oF)

- RDMA

- GPUDirect Storage

- Parallel file systems

- AI data pipelines

Networking & Infrastructure

- Strong networking knowledge:
- Bonding

- VLANs

- Routing

- MTU optimization

- DNS

- DHCP

- Experience with high-performance data center networking:
- 100G/200G/400G Ethernet

- RoCE

- RDMA

- Spine-Leaf architectures

- Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies

- Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting

Operations & Reliability

- Experience with high availability, clustering, and disaster recovery

- Strong troubleshooting skills across:
- Linux operating systems

- Hardware platforms

- GPU infrastructure

- Networking

- Enterprise storage

- Experience supporting mission-critical pro

← All remote jobs

Get remote developer jobs like this by email

One weekly digest. No spam, unsubscribe anytime.

Similar for you