Senior Site Reliability Engineer
Remote
WorldwideremotePosted 14 days ago
Skill Required
Site-Reliability-EngineeringPlatform-EngineeringDevOpsCloud-EngineerInfrastructure-EngineeringSenior-Site-Reliability-EngineerSenior-Site-Reliability-Engineering-ArchitectSenior-Reliability-EngineerSite-Reliability-Engineering-LeadSite-Reliability-Operations-EngineerSite-Reliability-Engineer-IIDevOps-Site-Reliability-EngineerSite Reliability EngineerSOCGitHub ActionsObservabilityEngineeringR ProgrammingKubernetesPrometheusautomationTerraformsimilar)securitybuildingEmbedded CDockerPythonGolangdesignNode.jsGitLinuxCI/CDCloudDesign PatternsShell ScriptingFulltime
Key highlights
- Salary: $53,300—$119,850 USD annually
- Level Required: Solid professional experience in SRE, DevOps, or Platform Engineering
- Key benefit: Fully remote with async work, 16 weeks paid parental leave, stock options, learning budget
- Notable requirement: Practical, embedded use of AI in infra/ops/dev work with agentic workflows
- Location priority: Europe for this hire due to diversity and timezone requirements
- Application process includes async infrastructure exercise (2-4 hours)
Role overview
Remote is a global employment compliance platform that enables businesses to recruit, pay, and manage international teams. This senior-level role focuses on reliability engineering and platform architecture in a fully async, remote-first environment across 6 continents. The company emphasizes innovation through AI and automation, with core values of trust, inclusion, and ambition. The SRE team operates with high autonomy on complex platform reliability challenges, collaborating closely with product and security teams.
Responsibilities
- Work with a high degree of autonomy on complex reliability and platform problems, owning the plan and execution of features and projects within the SRE/Platform domain
- Contribute to the platform's architecture and reliability strategy, translating ambiguous requirements into robust, maintainable solutions
- Raise the technical bar of engineers around you while collaborating closely with product and security teams in an async-first, fully remote environment
- Build reusable AI workflows that make the whole team faster and more reliable, not just yourself
- Lead solution discovery and delivery for reliability and infrastructure problems with real ambiguity, complexity, or scope
- Contribute to the platform's architecture, tooling, and roadmap; influence team priorities and advocate for technical initiatives
- Help define and operate reliability practices for the platform: SLOs/SLIs, error budgets, alerting, observability; take responsibility for the team's operational stance, using support/incident metrics to shape technical strategy
- Resolve cross-team requests, identify systemic issues, and turn recurring ones into reusable fixes and runbooks rather than one-off answers
- Work AI-natively and operationalise it for the team: use agentic workflows by default; build reusable prompts, skills, and tooling embedded in the codebase so others ship faster, safely; design agent-ready systems (clean interfaces, good observability) that make AI-assisted changes easy to review; establish shared standards and domain-level guardrails (secure-by-default patterns, CI protections, AI-assisted review practices)
- Mentor and give timely, actionable feedback to less-senior engineers; participate in hiring, onboarding, and RFC discussions
- Collaborate with Security on platform hardening and threat mitigation; contribute to capacity and cost-efficiency of the infrastructure
- Participate in incident response and on-call rotations to rapidly resolve issues and maintain system reliability
- Report to: SRE Team Lead
- Work in the Engineering team
- Work fully remotely with async collaboration
Requirements
- Solid professional experience in SRE, DevOps, or Platform Engineering
- Solid hands-on Kubernetes: operating and scaling production clusters and container tooling (Docker) and its ecosystem
- Experience building and managing cloud infrastructure on AWS (or similar)
- Strong infrastructure-as-code practice with Terraform
- Experience with reliability frameworks: SLOs, SLIs, error budgets, alerting strategies
- Solid observability background: OpenTelemetry, Grafana/Prometheus or similar
- Proficiency with CI/CD (GitLab CI, GitHub Actions, or similar) and deployment automation
- Comfortable with Golang, Bash/scripting; broader programming a plus
- Practical, embedded use of AI in infra/ops/dev work, agentic workflows with concrete, observable results, not just familiarity with the tools
- Clear and thoughtful communication, especially in an async-first, global setting
- Proactive, curious, and comfortable taking ownership of challenges
- Collaborative and respectful across cultures, time zones, and backgrounds
Nice to have
- Experience with 1 back-end programming language (Elixir, Nodejs, Python, etc)
- Experience running and configuring Linux systems in a non-cloud environment
- Security knowledge and capabilities from a defensive and offensive standpoint
Benefits
- Work from anywhere
- Flexible paid time off
- Flexible working hours (async work)
- 16 weeks paid parental leave
- Mental health support services
- Stock options
- Learning budget
- Home office budget & IT equipment
- Budget for local in-person social events or co-working spaces
Additional details
- For this hire, due to diversity and timezones requirements, we're prioritising Europe
- Start date: As soon as possible
- Application process: Interview with recruiter, Interview with HM, (async) Infrastructure exercise (2-4 hours), Interview with the team (without manager), Bar Raiser Interview, Executive Interview, Offer + Background check (Veremark & Remote)
- Interview with recruiter
- Interview with HM
- Async infrastructure exercise (not expected to spend more than 2-4 hours)
- Interview with the team (without any manager in the call)
- Bar Raiser Interview
- Executive Interview
- Offer + Background check (Veremark & Remote)
- The annual salary range for this full-time position is $53,300—$119,850 USD
- Location-based compensation using geo ranges to consider geographic pay differentials
- Actual base pay dependent upon location, transferable or job-related skills, work experience, relevant training, business needs, and market demands
- Salary range may be subject to change
- Foster internal mobility as a key element of culture of employee growth and development
- Compensation philosophy guarantees pay equity and fairness
- All compensation changes associated with internal move reviewed by Total Rewards & People Enablement team on case by case basis
- Remote's Total Rewards philosophy ensures fair, unbiased compensation and fair equity pay along with competitive benefits
- Do not agree to or encourage cheap-labor practices; ensure to pay above in-location rates
- Hope to inspire other companies to support global talent-hiring and bring local wealth to developing countries
- At first glance salary bands seem quite wide with context about geo ranges and global compensation strategy
- All positions are fully remote
- Team works async around the world
- Core values: Innovation, Trust, Inclusion, Ambition, Future Focus
- Team works tirelessly on ambitious problems asynchronously
- Apply now and define the future of work
- Please fill out the form below and upload CV with PDF format
- Submit application and CV in English (standardised language)
- If no up-to-date CV, LinkedIn profile acceptable instead
- Voluntarily tell us your pronouns at interview stage
- Option to answer anonymous demographic questionnaire when applying
- Equal employment opportunity employer; workforce should reflect people of all backgrounds, identities, and experiences
- Please note we accept applications on an ongoing basis
- Originally posted on Himalayas