Lead Infrastructure Engineer - Remote
ClanX
WorldwideremotePosted 16 days ago
Skill Required
Infrastructure-EngineeringCloud-EngineerKubernetes-EngineeringPlatform-EngineeringSite-Reliability-EngineeringLead-Infrastructure-EngineerInfrastructure-Operations-LeadInfrastructure-LeadInfrastructure-Team-LeadSenior-Infrastructure-EngineerCloud-Infrastructure-LeadIT-Infrastructure-LeadInfrastructure-EngineerInfrastructure EngineertroubleshootingNginxVMwareObservabilityEngineeringKubernetesnetworkingautomationdesigningsecuritybuildingdesignLinuxDNSandFulltime
Key highlights
- Lead the infrastructure platform of a GPU-focused neocloud, spanning bare metal, Kubernetes, networking, storage, automation, and reliability.
Role overview
Lead Infrastructure Engineer with 8+ years of infrastructure experience to own Kubernetes and core infrastructure for a GPU-focused neocloud platform. The role involves owning Kubernetes and core infrastructure for a remote-first neocloud provider building a full-stack GPU-focused edge platform from bare metal to inference services.
Responsibilities
- Design, build, and operate production-grade Kubernetes infrastructure.
- Design virtualized Kubernetes clusters and tooling for provisioning, upgrades, scaling, security, observability, and troubleshooting.
- Build Infrastructure as Code and automation for repeatable provisioning and operations.
- Design and operate Kubernetes networking, ingress, DNS, service discovery, load balancing, and network policies.
- Design and operate Kubernetes storage, persistent volumes, and CSI-based systems.
- Diagnose complex infrastructure problems across multiple layers of the stack.
- Build internal platform services that improve reliability or developer experience.
- Establish infrastructure standards, operational practices, observability, and reliability mechanisms.
- Make pragmatic trade-offs between speed, reliability, simplicity, and maintainability.
- Provide technical leadership and shape the platform architecture.
- Experience operating containerized workloads in production.
- Ability to debug across Kubernetes, Linux, networking, storage, virtualization, and physical infrastructure.
- Experience building reliable, observable, and operable infrastructure.
- Ability to own infrastructure problems from architecture through production.
Requirements
- 9+ years of experience in infrastructure engineering.
- Deep understanding of Kubernetes internals and core components.
- Strong production experience designing and operating Kubernetes environments.
- Strong Infrastructure as Code and automation experience.
- Strong understanding of Kubernetes networking, ingress, DNS, service discovery, load balancing, and storage.
- Strong Linux and systems troubleshooting skills.
- Experience building reliable, observable, and operable infrastructure.
- Ability to own infrastructure problems from architecture through production.
- Strong technical leadership and pragmatic decision-making.
Benefits
- Remote
- Interview Process: Recruiter Screening, Infrastructure Deep Dive, Kubernetes & Systems Design, Technical Leadership Round, Founder/Team Round
Additional details
- NeoHPC is a small, remote-first neocloud provider building a full-stack GPU-focused edge platform from bare metal to inference services.
- Important Note: ClanX is a recruitment partner, helping NeoHPC hire Lead Infrastructure Engineer.
- Originally posted on Himalayas