Member of Technical Staff, Production Engineering Lead
Pure Storage
Bangalore, IndiaPosted 1 month ago
Skill Required
EngineeringTechnical LeadGenerative AIBackup and RecoveryObservabilityWeb AccessibilityPrometheusKubernetesRabbitMQsecurityPythonGolangDockerdesignOpenAPIKafkaRedisCloudAPIsRAGElasticsearchGoAI
Key highlights
- Required experience: 8+ years in Platform Infrastructure Engineering with a shift toward AI Application development.
- Must-have languages: strong proficiency in Python (for AI logic) and Go (for platform tools).
- Key benefit: Flexible time off, wellness resources, company-sponsored team events, and additional perks via purebenefits.com.
- Notable requirement: Deep experience with Kubernetes, specifically managing GPU workloads and specialized storage for vector databases, plus proven experience with Kafka or RabbitMQ.
- Workplace accolades: Named to Fortune's Best Workplaces in Technology™ and Fortune's Best Workplaces in the Bay Area™, and certified as a Great Place to Work®.
- Work arrangement: On-site (#LI-ONSITE).
Role overview
We are seeking an experienced Software Engineer with a strong background in Platform Engineering to join the AI Applications team. In this role, you will be responsible for the "Inner Loop" of AI development—ensuring that LLM-based applications, RAG pipelines, and agentic workflows are built, tested, and deployed with the same rigor as traditional software. You will bridge the gap between high-velocity AI experimentation and stable, enterprise-grade production infrastructure, working within a fundamentally industry-reshaping tech environment.
Responsibilities
- Create internal APIs and abstractions that allow Software Engineers to provision AI-ready environments (complete with model weights, vector DBs, and event streams) with a single command (Platform Self-Service).
- Use Golang and Python to build internal tools and AI Agents that automate root-cause analysis of infrastructure failures and proactively optimize infrastructure (Agentic Automation).
- Leverage Kafka or RabbitMQ to build asynchronous AI processing pipelines (e.g., long-running document ingestion for RAG systems) (Event-Driven AI Integration).
- Ensure "Golden Path" deployments for AI models using Docker and Kubernetes, ensuring that the model, the prompt, and the code are all perfectly synced across environments (Reproducibility).
- Design and implement high-performance Go services that listen to Kafka/RabbitMQ streams to trigger dynamic infrastructure scaling based on real-time AI model demand (Event-Driven Scaling).
- Manage and optimize the infrastructure for Vector Databases and distributed caches (Redis), ensuring high availability for RAG data (Distributed State & Storage).
- Replace brittle shell scripts with robust, type-safe Internal Tooling in Go for automated environment provisioning and disaster recovery (Go-Based Tooling).
- Using Prometheus and Grafana to monitor not just system health, but AI-specific metrics like token latency and model "drift" (Observability).
- Write clean, maintainable, and testable Go code to manage complex cloud environments, viewing every infrastructure problem as a software challenge ("Code-First" Infrastructure Mindset).
- Use Go's goroutines and channels to handle thousands of concurrent events, managing the massive data throughput required by Kafka-driven AI pipelines (High-Concurrency Systems).
- Extend the K8s API with custom controllers to make the cluster "AI-aware" (Deep Orchestration Knowledge).
- Translate the unique infrastructure needs of AI (GPU memory management, high-speed NVMe storage, and vector retrieval) into stable, scalable systems (Bridge Between Data & Ops).
- Ensure a "Zero-Downtime" philosophy through Monitoring (Grafana/Prometheus) and Log Management (ELK), ensuring that even under heavy AI inference loads, the system remains performant and observable (Operational Resilience & Reliability).
- Apply sophisticated understanding of asynchronous architecture, knowing when to use Kafka for high-volume streaming versus RabbitMQ for complex task routing (Event-Driven Strategy).
- Ensure that data privacy (vital for AI) is baked into the infrastructure layer through a "Security-as-Code" approach (Security & Compliance Guardrails).
- Bring a mindset of collaboration, reliability, and continuous improvement to strengthen team productivity and delivery speed.
Requirements
- 8+ years in Platform Infrastructure Engineering with a shift toward AI Application development.
- Strong proficiency in Python (for AI logic) and Go (for platform tools).
- Practical experience with RAG (Retrieval-Augmented Generation), Prompt Engineering, and integrating LLM APIs (OpenAI, Anthropic, or local models via Ollama) (AI Literacy).
- Deep experience with Kubernetes, specifically managing GPU workloads and specialized storage for vector databases.
- Proven experience with Kafka or RabbitMQ for managing high-volume data streams.
- Expertise in using Go's goroutines and channels to handle thousands of concurrent events for Kafka-driven AI pipelines.
- Deep orchestration knowledge with Kubernetes, including extending the K8s API with custom controllers to make clusters "AI-aware."
- Operational Resilience & Reliability experience with Monitoring (Grafana/Prometheus) and Log Management (ELK) and a "Zero-Downtime" philosophy.
- Sophisticated understanding of asynchronous architecture, knowing exactly when to use Kafka for high-volume streaming versus RabbitMQ for complex task routing.
- A "Security-as-Code" approach to data privacy in infrastructure.
Benefits
- Flexible time off to manage a healthy work-life balance.
- Wellness resources.
- Company-sponsored team events.
- Additional perks available via purebenefits.com.
- Named to Fortune's Best Workplaces in Technology™.
- Named to Fortune's Best Workplaces in the Bay Area™.
- Certified as a Great Place to Work®.
- Employee Resource Groups to cultivate community.
- Inclusive leadership advocacy.
- Equal opportunity employment — no discrimination based on race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or any other characteristic legally protected by the laws of the jurisdiction in which you are being considered for hire.
- Accommodations available for candidates with disabilities during all aspects of the hiring process (contact TA-Ops@purestorage.com if invited to an interview).
Additional details
- Company (Everpure/Pure Storage) is in an unbelievably exciting area of tech and is fundamentally reshaping the data storage industry.
- The company encourages innovative thinking and growth, inviting candidates to join "the smartest team in the industry."
- The type of work described changes the world and is what the tech industry was founded on.
- The role is on-site, as indicated by the #LI-ONSITE tag.
- The requisition/level tag referenced is #LI-KT7.
- Candidates are encouraged to "bring your bold" and join a team that celebrates those who think critically, like a challenge, and aspire to be trailblazers (Innovation).
- The company gives space and support to grow and contribute to something meaningful (Growth).
- The team builds each other up and sets aside ego for the greater good (Team).
- The company is committed to fostering the growth and development of every person.
- The company is forging a future where everyone finds their rightful place and where every voice matters; uniqueness is accepted and embraced.