Skip to main content
Zobhira
Home
Jobs
Certifications
Zobhira
JobsCertificationsTodayAbout
Log inSign up
Zobhira

New job and contest openings, updated every morning on one searchable board.

Find work

  • All jobs
  • Fresher roles
  • Remote roles
  • Certifications

Compete

  • Added today

Popular cities

  • India
  • Bangalore, Karnataka, India
  • Hyderabad, Telangana, India
  • Pune, Maharashtra, India
  • Chennai, Tamil Nadu, India
  • Mumbai, Maharashtra, India

Company

  • About
  • Contact
  • Privacy
  • Terms

Stay updated

One email a week with new roles.

Secure infrastructure
Free to use, no account needed to search

Board updated daily · © 2026 Zobhira. All rights reserved.

Privacy PolicyTerms of Service
Home / Jobs / uvation

Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

uvation

IndiaremotePosted 2 days ago
uvation logo

Skill Required

Linux-Infrastructure-EngineeringBare-Metal-InfrastructureAI-Factory-InfrastructureGPU-Infrastructure-EngineeringStorage-EngineeringLinux-Infrastructure-EngineerLinux-Infrastructure-SpecialistInfrastructure-Linux-Unix-AnalystSenior-AI-Infrastructure-EngineerAI-Infrastructure-EngineerAI-ML-Infrastructure-EngineerMachine-Learning-Infrastructure-EngineerInfrastructure-Platform-EngineerInfrastructure-EngineerAI EngineerInfrastructure EngineerLinuxAIQuery OptimizationBackup and RecoverytroubleshootingVMwareObservabilityEngineeringKubernetesPrometheusnetworkingautomationTerraformdesigningFirmwarebuildingAnsiblePythondesignDevOpsArgoCDAzureCI/CDCloudShell ScriptingContract

Key highlights

  • Senior Linux Infrastructure Engineer role (senior level).
  • Requires expert‑level Linux administration (Ubuntu required).
  • Must have experience with GPU platforms such as A100/H100/H200/B200.
  • Must have experience with BMaaS and enterprise storage technologies (e.g., Ceph, WEKA, VAST Data).
  • Role explicitly not DevOps‑focused; dedicated DevOps team already exists.
  • Candidate must support mission‑critical production environments.

Role overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

Responsibilities

  • Designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.
  • Creating operational documentation, runbooks, and infrastructure standards.
  • Supporting mission‑critical production environments.
  • Bash and Python scripting for automation and operational efficiency.

Requirements

  • Expert‑level Linux administration (Ubuntu required; Red Hat and SUSE preferred).
  • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management.
  • Experience operating Bare Metal as a Service (BMaaS) platforms and large‑scale infrastructure environments.
  • BIOS/UEFI.
  • RAID controllers.
  • Firmware management.
  • iLO/iDRAC/IPMI.
  • NICs and SmartNICs.
  • HBA cards.
  • Hardware diagnostics and troubleshooting.
  • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale.
  • Experience deploying and managing GPU‑accelerated infrastructure for AI/ML workloads.
  • A100, H100, H200, B200, or equivalent GPU platforms.
  • NVIDIA DGX and OEM GPU servers.
  • GPU provisioning and lifecycle management.
  • GPU monitoring and performance optimization.
  • Knowledge of AI Factory architecture and infrastructure requirements.
  • Experience supporting GPU clusters, AI training environments, and high‑performance computing (HPC) workloads.
  • GPU resource allocation and scheduling.
  • Multi‑GPU systems.
  • GPU networking requirements.
  • High‑bandwidth, low‑latency infrastructure design.
  • CUDA.
  • NCCL.
  • GPUDirect Storage.
  • NVIDIA Fabric Manager.
  • LVM.
  • XFS, EXT4.
  • NFS.
  • iSCSI.
  • Fibre Channel SAN.
  • Multipath I/O.
  • Cluster architecture.
  • MON, OSD, MDS.
  • RBD, CephFS, RGW.
  • Capacity planning.
  • Performance tuning.
  • Failure recovery.
  • WEKA.
  • VAST Data.
  • Dell PowerScale.
  • Pure Storage FlashBlade.
  • NetApp.
  • NVMe‑over‑Fabrics (NVMe‑oF).
  • RDMA.
  • Parallel file systems.
  • AI data pipelines.
  • Bonding.
  • VLANs.
  • Routing.
  • MTU optimization.
  • DNS.
  • DHCP.
  • 100G/200G/400G Ethernet.
  • RoCE.
  • Spine‑Leaf architectures.
  • Familiarity with NVIDIA Spectrum‑X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies.
  • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting.
  • Experience with high availability, clustering, and disaster recovery.
  • Linux operating systems.
  • Hardware platforms.
  • GPU infrastructure.
  • Networking.
  • Enterprise storage.

Nice to have

  • Kubernetes infrastructure (especially AI/ML and GPU integration).
  • KVM, VMware, OpenShift Virtualization, or similar virtualization platforms.
  • Ansible automation.
  • NVIDIA Base Command Manager.
  • Slurm or HPC workload schedulers.
  • Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry).
  • Data Center Infrastructure Management (DCIM) tools.
  • IPAM solutions.
  • AWS, Azure, or hybrid cloud exposure.
  • DevOps experience is a plus.

Additional details

  • This is not a DevOps‑focused role. We already have a dedicated DevOps team and are looking for an engineer with extensive hands‑on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high‑performance storage, data center operations, and enterprise Linux platforms.
  • The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI‑ready platforms.
  • They should be comfortable working with high‑performance computing (HPC), AI Factory environments, and large‑scale Linux deployments where performance, reliability, and operational excellence are critical.
  • We are not looking for candidates whose experience is primarily CI/CD pipeline engineering, engineers focused mainly on Terraform, GitOps, or application delivery pipelines, cloud‑only administrators with limited bare metal, storage, or hardware experience, or professionals whose primary expertise is software development rather than infrastructure engineering.
  • Ideal Candidate: Someone who has spent years designing, building, and operating enterprise Linux environments, large‑scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU‑enabled infrastructure, BMaaS platforms, enterprise storage, and high‑performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges.
  • Originally posted on Himalayas.

Similar jobs open now

C
onsite
DevOps Engineer
Cantellat Solutions
Hyderabad, Telangana, India
View details
Archimedis Digital logo
Data Scientist – Databricks (4 years of experience)
Archimedis Digital
Chennai, Tamil Nadu, India
View details
Nihilent logo
Senior Network Engineer
Nihilent
Pune, Maharashtra, India
View details
Nihilent logo
Senior Network Engineer
Nihilent
Chennai, Tamil Nadu, India
View details
Apply now
LocationIndia
TypeContract
Posted9/8/2026
Apply by11/7/2026

Links are checked every day. If this one stops working, tell us and we'll pull it.

More like this

View all
C
DevOps Engineer
Cantellat Solutions
Archimedis Digital logo
Data Scientist – Databricks (4 years of experience)
Archimedis Digital
Nihilent logo
Senior Network Engineer
Nihilent
Apply now