Staff Site Reliability Engineer
Harvey
Bengaluru, IndiaPosted 5 months ago
Skill Required
EngineeringSite Reliability EngineerSOCDockerCloud SecurityAWSObservabilityR ProgrammingKubernetesnetworkingautomationTerraformsecuritybuildingDatadogwrittenetc.)PythondesignAzureCI/CDCloudDesign PatternsShell ScriptingGCPandFulltime
Key highlights
- Based in San Francisco, CA (in-person work model)
- 12+ years of SRE experience required
- Relocation assistance offered
- Must be authorized to work in India (no visa sponsorship)
Role overview
At Harvey, we’re transforming how legal and professional services operate by combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise. As a Staff Software Engineer on the Site Reliability (SRE) team, you will ensure the reliability, scalability, and performance of our legal AI platform. You’ll join a high-leverage team at the intersection of infrastructure and product, owning the systems that keep our platform fast, secure, and always on. From scaling across 50+ regions to automating mission-critical operations, your work will ensure Harvey remains resilient as we grow.
Responsibilities
- Design, implement, and manage monitoring, alerting, and infrastructure resources (compute, storage, networking) across 50+ global regions
- Lead incident management processes, including postmortems, root cause analyses, and driving actionable improvements
- Automate operational tasks and workflows, building tools and processes for capacity planning, graceful rollouts, and safe data access to maintain high reliability and reduce manual intervention
- Establish best practices for security, compliance, and reliability and collaborate across teams to drive these principles throughout the software lifecycle
- Optimize infrastructure costs through strategic capacity planning and build-versus-buy decisions while maintaining system performance, reliability, and functionality
- Provide technical mentorship and leadership, promoting best practices and fostering team growth
Requirements
- 12+ years of experience in Site Reliability Engineering (SRE) or similar roles supporting production environments, with proven ability to mentor and guide technical teams
- Expertise in infrastructure as code (IaC) tools (Pulumi, Terraform, CloudFormation, etc.)
- Deep familiarity with observability tools (Datadog, Sentry, etc.) and incident response practices (PagerDuty, IncidentIO, etc.)
- Proficiency with cloud infrastructure platforms (Azure, GCP, AWS, etc.)
- Strong programming skills (Python, Bash, Go, or similar languages)
- Proven track record of diagnosing complex system problems and implementing durable solutions
- Solid understanding of CI/CD, Kubernetes, containerization, networking, databases, and cloud security principles
- Excellent problem-solving skills, meticulous attention to detail, and a commitment to operational excellence
- Must be authorized to work in India. Visa sponsorship is not available for this role
Benefits
- Relocation assistance to new employees
- Commitment to providing reasonable accommodations to applicants with disabilities
Additional details
- This role is based in San Francisco, CA. We use an in-person work model
- Please find the Applicant Privacy for this region here
- #LI-AS2
- Harvey is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law
- Requests for accommodations can be made by emailing accommodations@harvey.ai