Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

uvation

AI/MLPlatform/InfraEmbedded

Apply on the company’s site

Frontier is not the employer and does not collect applications.

About this role

Job Overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms .

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills

Linux & Bare Metal Infrastructure

• Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)

• Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management

• Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments

• Strong understanding of server hardware, including:

• BIOS/UEFI

• RAID controllers

• Firmware management

• iLO/iDRAC/IPMI

• NICs and SmartNICs

• HBA cards

• Hardware diagnostics and troubleshooting

• Experience designing, implementing, and supporting enterprise Linux infrastructure at scale

AI Factory & GPU Infrastructure

• Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads

• Understanding of NVIDIA GPU technologies including:

• A100, H100, H200, B200, or equivalent GPU platforms

• NVIDIA DGX and OEM GPU servers

• GPU provisioning and lifecycle management

• GPU monitoring and performance optimization

• Knowledge of AI Factory architecture and infrastructure requirements

• Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads

• Understanding of:

• GPU resource allocation and scheduling

• Multi-GPU systems

• GPU networking requirements

• High-bandwidth, low-latency infrastructure design

• Familiarity with NVIDIA ecosystem technologies such as:

• CUDA

• NCCL

• GPUDirect Storage

• NVIDIA Fabric Manager

• NVIDIA Base Command (preferred)

Enterprise Storage & Data Platforms

• Advanced Linux storage administration:

• LVM

• XFS, EXT4

• NFS

• iSCSI

• Fibre Channel SAN

• Multipath I/O

• Strong hands-on experience with Ceph , including:

• Cluster architecture

• MON, OSD, MDS

• RBD, CephFS, RGW

• Capacity planning

• Performance tuning

• Failure recovery

• Experience with high-performance AI storage platforms such as:

• WEKA

• VAST Data

• Dell PowerScale

• Pure Storage FlashBlade

• NetApp

• Understanding of:

• NVMe-over-Fabrics (NVMe-oF)

• RDMA

• GPUDirect Storage

• Parallel file systems

• AI data pipelines

Networking & Infrastructure

• Strong networking knowledge:

• Bonding

• VLANs

• Routing

• MTU optimization

• DNS

• DHCP

• Experience with high-performance data center networking:

• 100G/200G/400G Ethernet

• RoCE

• RDMA

• Spine-Leaf architectures

• Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies

• Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting

Operations & Reliability

• Experience with high availability, clustering, and disaster recovery

• Strong troubleshooting skills across:

• Linux operating systems

• Hardware platforms

• GPU infrastructure

• Networking

• Enterprise storage

• Experience supporting mission-critical production environments

• Bash and Python scripting for automation and operational efficiency

• Experience creating operational documentation, runbooks, and infrastructure standards

Nice to Have

• Kubernetes infrastructure (especially AI/ML and GPU integration)

• KVM, VMware, OpenShift Virtualization, or similar virtualization platforms

• Ansible automation

• NVIDIA Base Command Manager

• Slurm or HPC workload schedulers

• Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)

• Data Center Infrastructure Management (DCIM) tools

• IPAM solutions

• AWS, Azure, or hybrid cloud exposure

We Are Not Looking For

• Candidates whose experience is primarily CI/CD pipeline engineering

• Engineers focused mainly on Terraform, GitOps, or application delivery pipelines

• Cloud-only administrators with limited bare metal, storage, or hardware experience

• Professionals whose primary expertise is software development rather than infrastructure engineering

Ideal Candidate

Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.

Originally posted on Himalayas

More at uvation

  • Software Engineer Frontend - React

    uvation

    Job Overview : As a React Frontend Software Engineer, you will be responsible for developing and implementing user interface components using React.js and workflows (such as Flux or Redux...

    Full StackFrontend

  • Network Security Engineer

    uvation

    Job Overview We are looking for an experienced Network Security Engineer to design, implement, monitor, and support enterprise security infrastructure across on-premises, cloud, and hybri...

    Cyber Security

  • Network Security Engineer

    uvation

    Job Overview We are looking for an experienced Network Security Engineer to design, implement, monitor, and support enterprise security infrastructure across on-premises, cloud, and hybri...

    Cyber Security

  • Network Security Engineer

    uvation

    Job Overview We are looking for an experienced Network Security Engineer to design, implement, monitor, and support enterprise security infrastructure across on-premises, cloud, and hybri...

    Cyber Security

  • Network Security Engineer

    uvation

    Job Overview We are looking for an experienced Network Security Engineer to design, implement, monitor, and support enterprise security infrastructure across on-premises, cloud, and hybri...

    Cyber Security

  • UI UX Designer

    uvation

    Job Title: UIUX Designer Job Description: The UIUX Designer is responsible the entire process of defining requirements, visualizing and creating graphics including illustrations, logos, l...

    Frontend