Staff Site Reliability Engineer

Harvey

Bengaluru, IndiaPosted 5 months ago
Harvey logo

Skill Required

EngineeringSite Reliability EngineerSOCDockerCloud SecurityAWSObservabilityR ProgrammingKubernetesnetworkingautomationTerraformsecuritybuildingDatadogwrittenetc.)PythondesignAzureCI/CDCloudDesign PatternsShell ScriptingGCPandFulltime

Key highlights

  • Based in San Francisco, CA (in-person work model)
  • 12+ years of SRE experience required
  • Relocation assistance offered
  • Must be authorized to work in India (no visa sponsorship)

Role overview

At Harvey, we’re transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. As a Staff Software Engineer on the Site Reliability (SRE) team, you will ensure the reliability, scalability, and performance of our legal AI platform. You’ll join a high-leverage team at the intersection of infrastructure and product, owning the systems that keep our platform fast, secure, and always on. From scaling across 50+ regions to automating mission-critical operations, your work will ensure Harvey remains resilient as we grow.

Responsibilities

  • Design, implement, and manage monitoring, alerting, and infrastructure resources (compute, storage, networking) across 50+ global regions
  • Lead incident management processes, including postmortems, root cause analyses, and driving actionable improvements
  • Automate operational tasks and workflows, building tools and processes for capacity planning, graceful rollouts, and safe data access to maintain high reliability and reduce manual intervention
  • Establish best practices for security, compliance, and reliability and collaborate across teams to drive these principles throughout the software lifecycle
  • Optimize infrastructure costs through strategic capacity planning and build-versus-buy decisions while maintaining system performance, reliability, and functionality
  • Provide technical mentorship and leadership, promoting best practices and fostering team growth

Requirements

  • 12+ years of experience in Site Reliability Engineering (SRE) or similar roles supporting production environments, with proven ability to mentor and guide technical teams
  • Expertise in infrastructure as code (IaC) tools (Pulumi, Terraform, CloudFormation, etc.)
  • Deep familiarity with observability tools (Datadog, Sentry, etc.) and incident response practices (PagerDuty, IncidentIO, etc.)
  • Proficiency with cloud infrastructure platforms (Azure, GCP, AWS, etc.)
  • Strong programming skills (Python, Bash, Go, or similar languages)
  • Proven track record of diagnosing complex system problems and implementing durable solutions
  • Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles
  • Excellent problem-solving skills, meticulous attention to detail, and a commitment to operational excellence
  • Must be authorized to work in India. Visa sponsorship is not available for this role

Benefits

  • Relocation assistance to new employees
  • Commitment to providing reasonable accommodations to applicants with disabilities

Additional details

  • This role is based in San Francisco, CA. We use an in-person work model
  • Please find the Applicant Privacy for this region here
  • #LI-AS2
  • Harvey is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law
  • Requests for accommodations can be made by emailing accommodations@harvey.ai
Apply now