Manager, System Software Engineering - Local AI
NVIDIA
India, PunePosted 1 month ago
Skill Required
AI EngineerEngineering ManagerSoftware EngineerSoftware DeveloperEngineeringAIMachine LearningData StructuresDeep LearningGenerative AIrelated fieldBusiness AnalysisR ProgramminganalyticaldebuggingbuildingPyTorchwrittenC++APIsFulltime
Key highlights
- 5+ years of industry experience required
- 2+ years of engineering leadership experience required
- Includes a generous benefits package
- Focuses on RTX, RTX Pro, and DGX-class systems
Role overview
NVIDIA is seeking a System Software Manager for the Local AI team to lead the development of an efficient on-device AI software stack for RTX, RTX Pro, and DGX-class systems. This role focuses on high-performance local inference, agentic workloads, low latency, efficient memory use, scalable infrastructure, and practical deployment on resource-constrained platforms to provide a streamlined experience for developers and end users.
Responsibilities
- Lead and grow a team building the on-device AI inference platform for RTX, RTX Pro, and DGX GPUs, with accountability for execution, technical direction, delivery quality, and roadmap alignment.
- Drive cross-functional alignment with NVIDIA’s software, research, architecture, and product teams, along with industry partners and open-source communities, to build strategy and strengthen the AI ecosystem across RTX and DGX platforms.
- Provide technical leadership for the architecture and evolution of modern inference runtimes and execution stacks across frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX, spanning workloads including LLMs, vision-language models, TTS, ASR, and diffusion models.
- Mentor engineers, develop technical leaders, and foster a high-performance team culture centred on innovation, collaboration, and operational excellence.
- Coordinate end-to-end optimization of AI models, data pipelines, and inference runtimes to improve performance across current and next-generation GPU architectures.
- Drive adoption of model optimization techniques such as quantization, pruning, sparsity, and distillation to enable efficient deployment of large models on local and edge devices.
- Establish team processes for system-level debugging, performance optimization, and performance-accuracy trade-off analysis, including infrastructure for performance and accuracy sweeps, gap analysis, and production-readiness improvements.
Requirements
- 5+ overall years of industry experience and 2+ years of engineering leadership experience.
- Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or a related field.
- Proven experience leading high-performing engineering teams in systems software, AI infrastructure, inference runtimes, or related domains.
- Strong technical foundation in C++ software development, debugging, data structures, algorithms, and machine learning systems.
- Extensive background in AI inference pipelines and Deep Learning frameworks like Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT.
- Deep understanding of inference backends and runtime internals, including scheduling, memory management, KV-cache behavior, graph execution, quantization, and hardware-aware optimization.
- Strong analytical and problem-solving skills, with the ability to balance technical depth, execution speed, and organizational priorities in a fast-paced environment.
- Excellent written and verbal communication skills, with proven ability to collaborate across engineering, product, research, and executive collaborators.
Nice to have
- Strong understanding of modern machine learning, deep neural networks, and generative AI, along with contributions to notable open-source projects.
- Demonstrated success building teams, setting technical vision, and scaling execution through periods of rapid growth.
- Track record of delivering end-to-end products with geographically distributed teams in multinational product organizations.
- Experience in lower-level systems or GPU programming, including CUDA and high-performance systems development.
- Contributions to open-source inference runtimes, model tooling, or performance infrastructure as well as practical experience working with frameworks and APIs including Llama.cpp, PyTorch, TensorRT, Vulkan, DirectX, and vLLM.
Benefits
- Competitive salaries.
- Generous benefits package.
Additional details
- NVIDIA has continuously reinvented itself for over two decades; the invention of the GPU in 1999 propelled the growth of PC gaming, redefined modern computer graphics, and revolutionized parallel computing.
- Recent focus on delivering AI models locally reduces latency, improves real-time processing, and addresses privacy concerns by minimizing data transfer to centralized servers.
- NVIDIA is an equal-opportunity employer and values diversity.
- The company is experiencing unprecedented growth, leading to the rapid expansion of engineering teams.