Kayzen is a mobile demand-side platform (DSP) dedicated to democratizing programmatic advertising, enabling leading apps, agencies, and brands to run high-scale campaigns with performance, transparency, and control. The Platform Engineering team owns complex, large-scale distributed systems like the Real-time bidding (RTB) platform, budget systems, and data pipelines, handling ~2M requests/sec with sub-millisecond latency and petabytes of data. The Lead DevOps Engineer will build and scale the backbone of Kayzen’s global RTB platform, engineering automated, self-healing infrastructure to treat private data centers like a programmable cloud.
Responsibilities
- Develop and maintain automated provisioning pipelines (PXE, ZTP) to deploy bare-metal servers (configure hardware, peripherals, services, settings, directories, storage, etc.) at scale across global data centers in accordance with standards and project/operational requirements
- Research and recommend innovative and automated approaches for system administration tasks
- Perform regular security monitoring to identify any possible intrusions
- Repair and recover from hardware or software failures, coordinating with impacted teams
- Apply OS patches and upgrades on a regular basis, and upgrade administrative tools and utilities
- Maintain data center environmental and monitoring equipment
- Perform ongoing performance tuning, hardware upgrades, and resource optimization as required
- Act as a technical lead and trusted point of contact for the infrastructure team, helping drive operational excellence and engineering best practices
- Support mentoring and onboarding of engineers, improve team collaboration and communication, and contribute to scaling the team as the infrastructure organization grows
- Report directly to CTO while partnering closely with other tech teams to improve reliability, automation, monitoring, and incident response processes across the infrastructure stack
Requirements
- Minimum 8 years of DevOps, system administration/debugging experience, scripting and related tools experience
- 2+ years of Team Lead/people management experience
- Flexibility to work in rotational shifts as part of the incident response team
- Good knowledge of coding with Shell and Python/Java and SQL commands
- Hands-on experience with Terraform or Ansible or Puppet/Chef for managing bare-metal configurations and automations
- Strong understanding of L4/L7 load balancing (HAProxy/Nginx) and network performance tuning (TCP/IP stack optimization)
- Good knowledge of pipeline/orchestration tools like Jenkins/Airflow and similar platforms
- Good knowledge of Observability platform experience (Prometheus, Grafana, InfluxDB)
- Hands-on experience with Unix/Linux
- Bachelor’s degree with a technical major, such as engineering or computer science
Nice to have
- Good knowledge of Kubernetes
- Systems Administration/System Engineer certification in Unix
Benefits
- Reporting directly to Co-founder & CTO
- Direct access to top management and an extremely “visible” role
- Exceptional career growth and learning opportunity
- Fully remote work setup
- Opportunity to be part of an experienced team of industry experts and entrepreneurs bringing massive change to the Adtech market
- Fun, driven, and multinational team located across multiple countries
- Flexible work-from-home arrangement
- 500-dollar home-office setup budget
- 1000-dollar annual learning and development budget
Additional details
- Location: Bangalore or Fully Remote from India
- Originally posted on Himalayas