Posted today · be early
AWS Cloud Engineer IV
Rackspace US, Inc.
IndiaremotePosted today
Skill Required
Cloud-EngineerDevOps-EngineerAWS-Cloud-EngineerInfrastructure-EngineeringCloud-ArchitectureAWS-EngineerSenior-AWS-EngineerAWS-Platform-EngineerAWS-DevOps-EngineerAWS-Systems-EngineerAWS-Infrastructure-EngineerAWS-Solutions-EngineerAWS-Managed-Services-EngineerCloud EngineerCloudAWSQuery OptimizationBackup and RecoverySOCGitHub ActionsWindows Serverrelated fieldObservabilityCD pipelinesEngineeringKubernetesPrometheusnetworkinganalyticalautomationShell ScriptingCryptographyTerraformsimilar)securitybuildingPythonAnsibleJenkinsFulltime
Key highlights
- 10+ years of relevant experience in cloud engineering, infrastructure, DevOps, or a related field is required
- The role acts as a technical SME focused on AWS, Terraform (IaC), operating systems, Kubernetes, and modern DevOps practices
- Key responsibilities include leading incident response, disaster recovery planning and testing, and mentoring L1/L2 engineers
- Preferred AWS certifications include Solutions Architect and DevOps Engineer Professional; beneficial certs include CNCF Kubernetes (CKA, CKAD, CKS) and Azure/GCP certs
- The role requires expert-level hands-on skills across cloud infrastructure, IaC, Kubernetes/containers, CI/CD, configuration management, operating systems, monitoring/observability, security/IAM, and disaster recovery
- Rackspace Technology is a multicloud solutions expert recognized as a best place to work by Fortune, Forbes, and Glassdoor
Role overview
The L4 Cloud DevOps Engineer serves as a technical subject matter expert (SME) for cloud architecture and operational engineering, with core expertise in AWS, Terraform (Infrastructure as Code), operating systems, Kubernetes (EKS), and modern DevOps practices. The role is responsible for designing, implementing, and supporting cloud infrastructure and deployment automation, leading incident response and disaster recovery activities, mentoring junior engineers, and collaborating with cross-functional stakeholders to deliver secure, reliable, and scalable solutions that align with business objectives.
Responsibilities
- Act as the L3 escalation point and technical SME for complex AWS, OS, Kubernetes, infrastructure and deployment issues
- Design and implement infrastructure as code using Terraform (and optionally CloudFormation), building reusable modules and enforcing coding standards
- Implement and maintain CI/CD pipelines and GitOps practices (ArgoCD), including pipeline design, environment promotion, rollback, and drift remediation
- Develop automation for configuration management and operational tasks using Ansible/AWX, scripting (Bash, PowerShell, Python), and other automation tools
- Design, deploy and maintain Kubernetes workloads (Amazon EKS), Helm charts, and environment-specific templating
- Perform Linux and Windows administration: hardening, patching, performance tuning, and automated deployments
- Lead incident response and major production incidents; perform root cause analysis (RCA) and implement permanent automated fixes
- Lead and coordinate Disaster Recovery planning and testing, including runbooks, failover/failback, and validation
- Define and enforce technical standards for IaC, Kubernetes, CI/CD, monitoring, logging, and operational processes
- Improve observability: monitoring, alerting, logging and operational reliability across platforms
- Mentor and provide technical guidance to L1 and L2 engineers; review and approve infrastructure and deployment changes, elevating team capability
- Create and maintain SOPs, runbooks, architecture documentation, and DR documentation
- Collaborate with product and customer engineering teams to deliver cloud-based solutions, migrations, and optimizations
- Lead and own complex engineering changes and transformations across cloud platforms
- Drive automation, standardization, and reliability improvements in operations and delivery
- Define technical standards, review peer code, and enforce best engineering practices
Requirements
- 10+ years of relevant experience in cloud engineering, infrastructure, DevOps, or a related field
- Deep technical expertise across AWS, Terraform, OS administration, Kubernetes, and DevOps tooling
- Ability to lead others to solve complex technical problems and provide technical leadership on projects
- Ability to work independently on sophisticated engineering tasks, escalating or seeking guidance only for the most complex, ambiguous situations
- Ability to provide functional leadership and mentorship to L1/L2 engineers
- Expert-level knowledge of AWS services, architecture patterns, security, networking, and operational best practices
- Deep understanding of Infrastructure as Code concepts and Terraform module design, state management and security
- Extensive knowledge of Linux and Windows internals, administration, automation, and security hardening
- In-depth Kubernetes architecture and operations experience
- Familiarity with common security and compliance frameworks (e.g., NIST, HIPAA, PCI) and secure operational controls
- Hands-on, expert-level skills in cloud and IaC: AWS, Terraform (including Terragrunt patterns), reusable module development, state management, and IaC security
- Hands-on, expert-level skills in Kubernetes and containers: design, operate, and troubleshoot EKS clusters, Helm chart development, container best practices, and GitOps (ArgoCD)
- Hands-on, expert-level skills in CI/CD and automation: build and maintain CI/CD pipelines, artifact management, automated testing, and deployment automation (Jenkins, GitHub Actions, GitLab CI, etc.)
- Hands-on, expert-level skills in configuration management: use Ansible/AWX for system configuration, patching, application deployment and repetitive operational tasks
- Hands-on, expert-level skills in operating systems: expert-level Linux administration, solid Windows server administration, and automation via scripting (Bash, PowerShell, Python)
- Strong Git skills (branching strategies, code review, hooks, access controls)
- Hands-on skills in monitoring and observability: implement monitoring, logging, and alerting (CloudWatch, Prometheus, ELK/EFK or similar)
- Hands-on skills in security and IAM: implement secure IAM practices, secret management, network security and encryption in cloud environments
- Hands-on skills in disaster recovery and resilience: DR planning, regular testing, and runbook creation for reliable failover and recovery
- Strong communication and stakeholder management skills, with the ability to translate technical concepts for non-technical audiences
- Analytical problem solving skills to diagnose complex production issues and identify root cause and long-term fixes
- Collaboration skills to work across disciplines to design, implement, and operate cloud solutions
- Project and time management skills to prioritize work, track deliverables, and meet deadlines while balancing multiple initiatives
- Technical documentation skills to create clear runbooks, SOPs, architecture diagrams, and standards
Nice to have
- Specific experience with Amazon EKS (preferred for Kubernetes operations role)
- Proficiency in CloudFormation (optional for infrastructure as code implementation)
- AWS Certified Solutions Architect or AWS Certified DevOps Engineer Professional certification (preferred)
- CNCF Kubernetes certifications (CKA, CKAD, CKS), Azure or GCP certifications (beneficial)
Additional details
- Rackspace Technology is a multicloud solutions provider that combines its expertise with the world's leading technologies across applications, data, and security to deliver end-to-end solutions, with a proven track record of advising customers based on their business challenges, designing scalable solutions, building and managing those solutions, and optimizing long-term returns.
- Rackspace Technology has been named a best place to work year after year by Fortune, Forbes, and Glassdoor, and focuses on attracting and developing world-class talent to fulfill its mission of embracing technology, empowering customers, and delivering the future.
- Rackers thrive through a shared central goal of being a valued member of a winning team on an inspiring mission, bring their whole selves to work every day, and embrace the notion that unique perspectives fuel innovation to enable the company to best serve its global customers and communities.
- Rackspace Technology is committed to offering equal employment opportunity without regard to age, color, disability, gender reassignment or identity or expression, genetic information, marital or civil partner status, pregnancy or maternity status, military or veteran status, nationality, ethnic or national origin, race, religion or belief, sexual orientation, or any legally protected characteristic.
- Applicants with a disability or special need that requires accommodation are invited to notify the company.
- This job posting was originally published on the Himalayas platform.