Site Reliability Engineer
HighRadius
Hyderabad, Telangana, IndiaPosted 1 month ago
Skill Required
Product & EngineeringSite Reliability EngineerBackup and RecoverySOCDockerCloud SecurityAWSVMwareObservabilityKubernetesPrometheusCryptographyTerraformAnsibleJenkinsPythonHadoopArgoCDLinuxMySQLAzureDesign PatternsITILDNSGCPGitElasticsearch
Key highlights
- 7+ years of industry experience required
- Expertise in AWS, Azure, or GCP required
- Valuation of $3.1B and $100M+ ARR
- Pre-IPO stage with rapid growth
- Recognized in Gartner Magic Quadrant and Forbes Cloud 100
- Global presence with 8+ locations
Role overview
HighRadius, a leading provider of cloud-based Autonomous Software for the Office of the CFO, seeks a highly skilled and adaptable Site Reliability Engineer (7+ years of experience) to join its Cloud Engineering team. In this role, you will design and refine cloud infrastructure with a focus on reliability, security, and scalability, applying software engineering principles to solve operational challenges and ensure system stability. The position involves managing live production environments, driving automation, and collaborating with cross-functional teams to integrate cloud solutions and improve resilience.
Responsibilities
- Design, build, and maintain resilient cloud infrastructure solutions to support the development and deployment of scalable and reliable applications.
- Manage and optimize cloud platforms for high availability, performance, and cost efficiency.
- Lead reliability best practices by establishing and managing monitoring and alerting systems to proactively detect and respond to anomalies and performance issues.
- Utilize SLI, SLO, and SLA concepts to measure and improve reliability.
- Identify and resolve potential bottlenecks and areas for enhancement.
- Contribute to the automation, provisioning, and standardization of infrastructure resources and system configurations.
- Identify and implement automation for repetitive tasks to significantly reduce operational overhead.
- Develop Standard Operating Procedures (SOPs) and automate workflows using tools like Rundeck or Jenkins.
- Participate in and help resolve major incidents, conduct thorough root cause analyses, and implement permanent solutions.
- Effectively manage incidents within the production environment using a systematic problem-solving approach.
- Work closely with diverse stakeholders and cross-functional teams, including software engineers, to integrate cloud solutions, gather requirements, and execute Proof of Concepts (POCs).
- Foster strong collaboration and communication.
- Guide designs and processes with a focus on resilience and minimizing manual effort.
- Promote the adoption of common tooling and components, and implement software and tools to enhance resilience and automate operations.
- Be open to adopting new tools and approaches as needed.
Requirements
- 7+ years of industry experience.
- Demonstrated expertise in at least one major cloud platform (AWS, Azure, or GCP).
- Extensive experience with containerization (Docker) and orchestration (Kubernetes) technologies.
- Proficiency in scripting languages (shell and Python).
- Experience with configuration management tools (Ansible or Puppet).
- Exposure to Infrastructure as Code (IaC) tools (Terraform or CloudFormation).
- Experience setting up and configuring monitoring tools (Prometheus, Grafana, or the ELK stack).
- Hands-on experience implementing OpenTelemetry for observability.
- Familiarity with monitoring and logging tools for cloud-based applications.
- Strong understanding of SLI, SLO, SLA, and error budgeting.
- Proven proficiency in on-premises hosting and virtualization platforms (VMware, Hyper-V, or KVM).
- Solid understanding of storage internals (NAS, SAN, EFS, NFS) and protocols (FTP, SFTP, SMTP, NTP, DNS, DHCP).
- Experience with networking and firewall technologies.
- Strong hands-on experience with Linux internals and operating systems (RHEL, CentOS, Rocky Linux).
- Experience with Windows operating systems supporting diverse environments.
- Excellent communication and interpersonal skills for effective teamwork.
- Proactive mindset with a willingness to learn and adapt in a dynamic environment.
- Pragmatic and adaptable mindset, with a willingness to step outside comfort zones and acquire new skills.
- Ability to consider the broader system impact of your work.
- Must be a change advocate for reliability initiatives.
Nice to have
- Experience with DevOps toolchain elements like Git, Jenkins, Rundeck, ArgoCD, or Crossplane.
- Experience with database management, particularly MySQL and Hadoop.
- Knowledge of cloud cost management and optimization strategies.
- Exposure to Gen AI.
- Understanding of cloud security best practices, including data encryption, access controls, and identity management.
- Experience implementing disaster recovery and business continuity plans.
- Familiarity with ITIL (Information Technology Infrastructure Library) processes.
Additional details
- HighRadius is a renowned provider of cloud-based Autonomous Software for the Office of the CFO, transforming critical financial processes for over 800+ leading companies worldwide.
- Trusted by prestigious organizations like 3M, Unilever, Anheuser-Busch InBev, Sanofi, Kellogg Company, Danone, Hershey's, and many others.
- HighRadius optimizes order-to-cash, treasury, and record-to-report processes.
- Recognized in Gartner's Magic Quadrant and a prestigious spot in Forbes Cloud 100 List for three consecutive years.
- Valuation of $3.1B and an impressive annual recurring revenue exceeding $100M.
- Robust year-over-year growth of 24%.
- Global presence spanning 8+ locations with a recent addition in Poland.
- In the pre-IPO stage, poised for rapid growth.
- Invites passionate and diverse individuals to join the path to becoming a publicly traded company.