Senior DevOps Engineer
DocuSketch
WorldwideremotePosted 25 days ago
Skill Required
DevOps-EngineerPlatform-EngineerSRECloud-Infrastructure-EngineerDevOps-LeadSenior-DevOpsSenior-DevOps-DeveloperSenior-Staff-DevOps-EngineerSenior-DevOps-Platform-EngineerSenior-AWS-DevOps-EngineerSenior-DevOps-ManagerSenior-DevOps-ArchitectDevOps EngineerDevOpsBackup and RecoverySOCGitHub Actionsrelated fieldObservabilityEngineeringR ProgrammingKubernetesautomationTerraformdesigningsecurityTestNGPythondesignArgoCDGitCI/CDCloudHelmShell ScriptingAWSandGoFulltime
Key highlights
- Requires 5+ years of professional experience in DevOps or related fields.
- Expert‑level experience with Kubernetes and Infrastructure‑as‑Code is mandatory.
- Strong knowledge of AWS multi‑account environments is required.
- Ability to use AI tools to improve reliability is a key requirement.
- Proven leadership of cross‑team infrastructure initiatives is essential.
- Professional fluency in English is mandatory.
Role overview
Responsibilities
- Lead complex infrastructure and delivery initiatives across teams and environments.
- Help evolve the platform toward self-service environments, clearer ownership of operations by product teams, and an SRE‑style operating model.
- Define and improve standards for cloud infrastructure, deployment, observability, and operational readiness.
- Design and operate scalable, secure, and cost‑efficient cloud platforms.
- Work with Infrastructure‑as‑Code, Kubernetes, containers, CI/CD, and automation frameworks.
- Improve monitoring, alerting, tracing, SLOs, disaster recovery, and incident‑management practices.
- Lead root‑cause analysis and drive sustainable post‑incident improvements.
- Build security and compliance into infrastructure and delivery processes, including least‑privilege access, identity federation and OIDC, short‑lived credentials, secrets management and rotation, auditability, certificate management, supply‑chain security, and policy‑as‑code.
- Implement automated testing, linting, validation, and rollback processes for infrastructure and delivery pipelines.
- Use AI agents responsibly to improve infrastructure design, automation, incident analysis, and operational workflows while applying strong engineering judgement and quality guardrails.
- Act as a technical partner to Engineering, Product, Security, and IT leadership.
- Mentor engineers, improve documentation and standards, and contribute to the long‑term direction of DevOps practices.
Requirements
- Strong professional experience in DevOps, platform engineering, infrastructure engineering, or a related field; typically 5+ years.
- Proven experience leading infrastructure or delivery initiatives beyond a single team or service.
- Strong knowledge of AWS or a comparable cloud platform, ideally including multi‑account environments.
- Expert‑level experience with Kubernetes, containers, Infrastructure‑as‑Code, and cloud automation.
- Hands‑on experience with Terraform, Helm, GitHub Actions, and GitOps tools such as Argo CD.
- Experience with production observability, including metrics, logs, traces, alerting, and platforms such as SigNoz.
- Experience designing reliable systems and improving operational readiness, observability, and incident response.
- Understanding of infrastructure security, compliance, identity, access management, and supply‑chain security.
- Experience creating automated testing, linting, validation, and rollback mechanisms for infrastructure or delivery pipelines.
- Strong problem‑solving skills and the ability to lead root‑cause analysis for complex incidents.
- Ability to communicate clearly with both technical and non‑technical stakeholders.
- Professional fluency in English.
- Ability to use AI tools effectively to deliver fast results and create reliable, repeatable workflows that improve the reliability and stability of systems.
Nice to have
- Experience with cost governance and cloud‑financial management.
- Experience with SLOs, error budgets, disaster recovery, and reliability engineering.
- Experience with policy‑as‑code and automated security scanning.
- Strong scripting or programming experience, for example Python, Go, or Bash.
- Experience introducing AI‑assisted engineering workflows with appropriate safeguards.
- Experience supporting AI‑platform or hybrid‑cloud workloads, including GPU scheduling or cluster operations with tools such as Karpenter, Cilium, or KubeRay.
- Experience mentoring engineers or contributing to organisation‑wide technical standards.
Additional details
- Originally posted on Himalayas.