Engineering Manager, SRE
Remote
Skill Required
Key highlights
- Annual salary range: $75,450—$169,700 USD
- Anytime start date
- 16 weeks paid parental leave
- Fully remote — work from anywhere in the world
- 4 direct reports reporting to Director of Engineering, Platform
- Role mix: 60% individual contributor, 40% leadership
Role overview
Remote is solving modern organizations' biggest challenge—navigating global employment compliantly with ease—helping businesses of all sizes recruit, pay, and manage international teams. With core values centered on Innovation and a future-focused, fully remote, async work culture, Remoters work across six continents. The SRE Team Leader role is a 60% individual contributor, 40% leadership position. You will own the career development of your reports, steer the team's focus against company goals, serve as the team's spokesperson across engineering, and stay close enough to technical work to set direction with credibility. The team owns Kubernetes, AWS, PostgreSQL, CI infrastructure, the observability stack, and the reliability practices built on top of it.
Responsibilities
- Own the career development of your reports
- Steer the team's focus using judgment against company goals
- Be the spokesperson for the team across engineering
- Stay close enough to technical work to set direction with credibility and to know when something is going wrong before it is escalated
- Manage the full career lifecycle of your reports: onboarding, feedback, performance assessment, progression, and hiring
- Coach both craft and soft skills
- Handle underperformance directly and early, with clarity and empathy
- Read team dynamics well and resolve conflict rather than routing around it
- Foster commitment to the goals of the company
- Prioritize exceptionally well when operational load and project work compete, and protect the team's focus without dropping operational commitment
- Write clearly for a fully distributed and async environment
- Build relationships across teams
- Define the SRE goals: what the team commits to, in what order, and why
- Manage the support rotation and on-call model
- Own Remote's core infrastructure: Kubernetes, AWS, PostgreSQL, DNS and TLS, CI infrastructure
- Own the reliability practice: SLOs, error budgets, incident response, and the observability stack
- Partner with the Security team on threats, patching, and infrastructure controls, including audit and compliance obligations
- Manage vendor relationships behind the platform, including renewals and commercial conversations with support from the Director
Requirements
- Hands-on background in site reliability, DevOps, or cloud infrastructure engineering, deep enough to review the team's work, challenge a design, and be taken seriously in an incident
- Kubernetes in production, including the operational reality of it rather than the happy path
- AWS at meaningful scale
- Hands-on AI building, enablement, and scaling AI infrastructure
- Solid observability practices and principles
- Infrastructure as code with Terraform
- CI/CD systems such as GitLab CI, GitHub Actions, or Jenkins
- Docker and shell scripting
- Experience running a reliability practice: incident response, on-call, SLOs, error budgets, and the discipline of turning incidents into changes that stick
- Understanding and history of working in regulated environments
- Proven people leadership: led an SRE, infrastructure, or platform engineering team and owned reports' growth, performance, and career progression rather than just their sprints
- Ability to tell the difference between a good interview and a good engineer
Nice to have
- Working knowledge of a backend language, ideally Elixir, or otherwise Java, Clojure, Node.js, Python, or similar
- Depth in modern observability: OpenTelemetry, distributed tracing, and tools such as Honeycomb
- Database operations experience, particularly PostgreSQL or Aurora performance, connection pool health, and query tuning
- Running and configuring Linux systems outside a cloud environment
- Security capability from both a defensive and an offensive standpoint
- Cloud cost management and FinOps
- Experience growing a team from a small base, including building the hiring bar as you go
Benefits
- Work from anywhere
- Flexible paid time off
- Flexible working hours (async)
- 16 weeks paid parental leave
- Mental health support services
- Stock options
- Learning budget
- Home office budget & IT equipment
- Budget for local in-person social events or co-working spaces
Additional details
- Team: Site Reliability Engineering, part of Platform Engineering
- Direct reports: 4
- Location: Anywhere in the world; current coverage strongest in EMEA and APAC, candidates who overlap with either or can help close the Americas gap are especially welcome
- Start date: As soon as possible
- Reports to: Director of Engineering, Platform
- Interview process: Interview with recruiter → Interview with hiring manager → Scenario interview with a peer team leader → Executive interview → Bar Raiser interview → Offer + Prior employment verification check
- Applications submitted in English; PDF CV required (LinkedIn profile accepted if no up-to-date CV available)
- Voluntarily share pronouns at interview stage; anonymous demographic questionnaire available at application
- Applications accepted on an ongoing basis
- Originally posted on Himalayas
- Remote uses geo ranges for geographic pay differentials as part of global compensation strategy
- Actual base pay depends on location, transferable/job-related skills, work experience, relevant training, business needs, and market demands
- Base salary range subject to change
- Internal mobility is fostered; compensation changes for internal moves reviewed by Total Rewards & People Enablement team case by case