Senior Site Reliability Engineer

Gradle

WorldwideremotePosted 1 month ago
Gradle logo

Skill Required

Site-Reliability-EngineeringDevOpsCloud-EngineerPlatform-EngineeringInfrastructure-EngineeringSenior-Site-Reliability-EngineerSenior-Site-Reliability-Engineering-ArchitectPrincipal-Site-Reliability-EngineerSenior-Reliability-EngineerSite-Reliability-Engineering-LeadSite-Reliability-Engineering-ManagerSite Reliability EngineerBackup and RecoverySOCObservabilityEngineeringKubernetesPrometheusTerraformbuildingwrittenPythonKotlinCloudTestNGJavaShell ScriptingAWSSAPGoAIFulltime

Key highlights

  • Remote from anywhere in Europe in the GMT timezone
  • 5+ years in SRE, DevOps, or equivalent role required
  • Competitive salaries and equity grants
  • Ground-floor role in a new SRE team
  • Strong Kubernetes and AWS expertise required
  • Remote-first work environment

Role overview

AI is changing how software gets built, shifting focus from writing code to orchestrating, verifying, and governing change. We build Develocity, a toolchain observability and intelligence platform used by leading organizations like Netflix, Airbnb, and Spotify. Develocity enables delivery excellence through deep observability, build/test acceleration, and AI-powered intelligence across the entire toolchain. As a founding member of our new SRE team, you will shape how we operate, ensuring the reliability, performance, and availability of Develocity instances and supporting infrastructure. You’ll work on our Cloud Application Platform, Kubernetes on AWS, and collaborate across teams to build reliability into our systems and processes.

Responsibilities

  • Operate and maintain all Develocity instances and supporting services.
  • Participate in a follow-the-sun on-call rotation, owning incident response and troubleshooting issues across the stack.
  • Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery.
  • Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
  • Work with engineering teams to build reliability into features from the start.
  • Run incident response and retrospectives, and make sure we learn from them.
  • Own disaster recovery, backups, and business continuity.
  • Communicate with customers during incidents and maintenance windows.
  • Optimize performance, resource usage, and costs.
  • Help evolve our SaaS operations as we grow.

Requirements

  • 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
  • Strong Kubernetes experience in production environments.
  • Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
  • Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
  • Track record of incident management and response.
  • Knowledge of SRE best practices (SLAs, SLOs).
  • Scripting proficiency (Python, Bash) for automation.
  • Experience with 24/7 on-call rotations.
  • Strong written and verbal English communication.

Nice to have

  • Experience operating SaaS platforms at scale.
  • Familiarity with Develocity.
  • JVM language experience (Java, Kotlin).
  • Disaster recovery planning and execution experience.
  • Customer-facing incident communication skills.
  • Experience establishing SRE practices in new or growing teams.

Benefits

  • A ground-floor role in a new SRE team—you'll shape how we do things, not inherit someone else's decisions.
  • Real ownership of production systems used by engineers at companies you've heard of.
  • Direct interaction with customers when things go wrong (and when they go right).
  • A culture that values automation over heroics.
  • In-person meetings, such as our annual company offsite and team meetings.
  • Work from home in a remote-first environment.
  • Competitive salaries and equity grants.

Additional details

  • We are an AI-native company. AI is not a feature we're bolting on – it's central to how we work, how we think about our product, and where we're heading.
  • We have partnered with the Apache Software Foundation, the Commonhaus Foundation, the Micronaut Foundation, and other OSS projects such as Spring, Quarkus, Kotlin, JUnit, AndroidX, and many more to bring the values of Develocity also to the OSS Community.
  • Our Values: Seek to Understand, Know the Why, Innovate & Iterate, Own the Outcome.
  • You'll be part of a distributed, remote-first team that values asynchronous communication and written documentation. Strong self-direction and clear communication across time zones are essential.
  • You'll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it.
  • Location: Remote from anywhere in Europe in the GMT timezone.
  • While our team works remotely and is spread across the globe, we deeply value daily interactions and collaboration.
Apply now