Senior DevOps Engineer
NVIDIA
Skill Required
Key highlights
- 5+ years of proven experience required
- Bachelor's or Master's Degree (or equivalent) in CS/Software Engineering required
- Strong Kubernetes and on-prem infrastructure expertise required
- Programming in Python/Golang/Java required
- Focus on AI-powered automation and infrastructure scaling
Role overview
NVIDIA is looking for an outstanding engineering lead to join its Software Infrastructure and Operations team. The position will be part of a fast-paced crew that develops and maintains sophisticated Kubernetes-based development, compute, and test environments for a multitude of platforms including Windows and Linux using OSS CICD tools like GitHub, GitLab, and Jenkins. You will work with a team of passionate and skilled engineers to provide better tools to build and manage this infrastructure, forging the next generation of compute infrastructure to multiply the power of CPU, GPU, and DPU for the age of AI. The role requires a motivated, hardworking, and focused individual with a real passion for operational excellence, infrastructure services, and automation.
Responsibilities
- Architect the scaling operation in our data centers
- Deploy and Support end-to-end container management solutions with Kubernetes, Docker, containerd
- Design solutions with service discovery, networking, monitoring, logging, scheduling in Kubernetes
- Manage end-to-end OSS CICD tools (GitLab/GitHub/Jenkins) in on-prem Kubernetes environment
- Design and develop tools needed for automating CICD & Developers workflow
- Design and build sophisticated automations and AI-powered applications
- Use depth in algorithms and system software background
- Work in teams to deploy new data center infrastructure
- Plan and implement critical metrics tracking using various data analytics mining methods and dashboards
- Reuse AI techniques to extract useful signals about machines and jobs from the data generated
- Take part in prototyping, crafting, and developing cloud infrastructure for NVIDIA
Requirements
- Strong Kubernetes understanding and background, especially on-premises setup and extensive experience with Kubernetes components & subsystems
- Experience maintaining large-scale on-prem infrastructure applications & OSS CICD tools using Kubernetes
- Proven programming background in Python/Golang/Java and/or relevant scripting languages
- Excellent debugging and analytical skills and experience in Databases (SQL: MySQL; NoSQL: Elastic Search/MongoDB)
- Proficient with configuration management tools like Ansible, Chef, Puppet
- Strong experience with Jenkins and/or other CI systems
- Hands-on experience with VMs, Docker, Kubernetes Cluster
- 5+ years of proven experience
- Bachelor's or Master's Degree or equivalent experience in CS, Software Engineering, or related field
Nice to have
- Experience with analytics/visualization tools like Kibana, Grafana, Splunk
- Experience with monitoring systems such as Zabbix and/or Nagios
- Previous experience with DevOps/SRE teams
- Thrives in a multi-tasking environment with constantly evolving priorities and documents work well
- Outstanding collaboration skills across organizational boundaries
- Experience with using and improving data centers and with computer algorithms, and ability to choose the best possible algorithms to meet the scaling challenge
- Ability to divide complex problems into simple sub-problems and then reuse available solutions to implement most of those
- Experience with designing simple systems that can work reliably without needing much support