Skip to main content
Zobhira
Home
Jobs
Certifications
Zobhira
JobsCertificationsTodayAbout
Log inSign up
Zobhira

New job and contest openings, updated every morning on one searchable board.

Find work

  • All jobs
  • Fresher roles
  • Remote roles
  • Certifications

Compete

  • Added today

Popular cities

  • India
  • Bangalore, Karnataka, India
  • Hyderabad, Telangana, India
  • Pune, Maharashtra, India
  • Chennai, Tamil Nadu, India
  • Mumbai, Maharashtra, India

Company

  • About
  • Contact
  • Privacy
  • Terms

Stay updated

One email a week with new roles.

Secure infrastructure
Free to use, no account needed to search

Board updated daily · © 2026 Zobhira. All rights reserved.

Privacy PolicyTerms of Service
Home / Jobs / Zeta Global

Lead Site Reliability Engineer

Zeta Global

Bengaluru, Karnataka, IndiaPosted 1 month ago
Zeta Global logo

Skill Required

EngineeringSite Reliability EngineerSystem DesignJMeterSOCShell ScriptingObservabilityCD pipelinesR ProgrammingKubernetesPrometheusnetworkingTerraformbuildingEmbedded CTestNGPythondesignLinuxCI/CDCloudDesign PatternsAWSElasticsearchGoCI

Key highlights

  • Position: Lead Site Reliability Engineer
  • Experience: 3-5 years of SRE experience
  • Key Tech: AWS, Kubernetes, and Infrastructure as Code (Terraform/Pulumi)
  • Programming: Strong skills in Python or Go
  • Notable Requirement: Participation in on-call rotation
  • Domain: Cloud-based and on-prem environments

Role overview

Zeta Global is a leading AI-powered marketing technology company that enables enterprise brands to acquire, grow, and retain customers through intelligent, data-driven engagement. As a Lead Site Reliability Engineer, the candidate will join a company founded in 2007 that unifies customer data, identity, and advanced analytics to transform billions of data signals into actionable marketing intelligence via the Zeta Marketing Platform (ZMP) and Athena by Zeta™.

Responsibilities

  • Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.
  • Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.
  • Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.
  • Automate incident detection and response using automated runbooks or predefined workflows.
  • Write software as needed to support reliability or efficiency needs.
  • Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.
  • Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.
  • Collaborate with development and operations teams on building reliable, scalable, and high-performance services.
  • Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.
  • Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.
  • Get involved in chaos engineering initiatives.
  • Participate on our on-call rotation.
  • Drive advanced alerting and anomaly detection applied to metrics

Requirements

  • 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.
  • Deep understanding of Linux systems, networking, and systems administration.
  • Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.
  • Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.
  • Strong skills in at least one programming language (Python, Go) to write production level code.
  • Strong skills in shell scripting using bash or similar.
  • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.
  • Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc.
  • Reliability-focused mindset with the ability to balance fast product iterations and system stability.
  • Solid understanding of SLOs, SLIs, and error budgets.
  • Hands-on knowledge of CI/CD pipelines and infrastructure automation.
  • Proven expertise in incident management, postmortems, and root cause analysis.
  • Knowledge of modern deployment strategies (e.g., blue-green deployments, canary releases) and resiliency patterns (circuit breakers, retry mechanisms, etc).

Nice to have

  • Experience with distributed systems.
  • Experience with statistical analysis applied to metrics.
  • Familiarity with high-performance, low-latency systems.
  • Strong problem solving skills.
  • Experience as on-call engineer.
  • Experience writing production level code.
  • Hands-on experience running Chaos Engineering drills and initiatives.

Additional details

  • Zeta Global was founded in 2007 by entrepreneur David A. Steinberg and John Sculley, former CEO of Apple Inc and Pepsi-Cola.
  • The Company combines the industry’s 3rd largest proprietary data set (2.4B+ identities) with Artificial Intelligence.
  • Publicly traded on the New York Stock Exchange (NYSE: ZETA).
  • Zeta redefined modern marketing with Athena by Zeta™, a superintelligent, conversational agent embedded in ZMP.

Similar jobs open now

ABB logo
Senior Project Engineer -SIS
ABB
Bangalore, Karnataka, India
View details
IBM logo
Data Engineer-Data Platforms-AWS
IBM
Pune, Maharashtra, India
View details
IBM logo
Technical Consultant-AI Integration
IBM
Pune, Maharashtra, India
View details
S
remote
AI Trainer Generalist
Sporterer
Mumbai, Maharashtra, India
View details
Apply now
LocationBengaluru, Karnataka, India
TypeNot specified
Posted7/15/2026
Apply byOpen

Links are checked every day. If this one stops working, tell us and we'll pull it.

More like this

View all
ABB logo
Senior Project Engineer -SIS
ABB
IBM logo
Data Engineer-Data Platforms-AWS
IBM
IBM logo
Technical Consultant-AI Integration
IBM
Apply now