DevOps Engineer
KnowledgeWorks Global Ltd.
Role tags
Tech stack mentioned
Role overview
Formatting this description...
Sr. DevOps Engineer Observability Strategy (Web & AI Applications) Design and implement an end-to-end observability & alert stack covering both traditional web applications and AI/ML services. Develop and maintain tooling such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalent Build AI-specific observability: model latency/throughput tracking, GPU utilization, token usage, drift detection, and inference quality signals. Infrastructure Scaling & Reliability: Design and manage infrastructure capable of scaling across on-premise data centers and public cloud Own capacity planning, load testing, auto-scaling, and cost-optimization initiatives across compute, storage, and networking. Implement Infrastructure as Code (Terraform, Ansible, or equivalent) to ensure environments are reproducible, version-controlled, and auditable. Lead disaster recovery, backup, and high-availability strategy for critical systems. DevOps & CI/CD Delivery : Partner closely with Solution Architects to translate project and system designs into concrete DevOps execution plans. Design, build, and maintain CI/CD pipelines (Jenkins, GitHub Actions or ArgoCD/Flux for GitOps) across multiple projects and teams. Containerize and orchestrate applications using Docker and Kubernetes, including Helm chart and manifest management. Embed security and compliance checks (SAST/DAST, secrets scanning, image scanning) directly into the delivery pipeline (DevSecOps). GPU Infrastructure & Model Deployment Provision, configure, and manage GPU infrastructure (on-prem clusters and cloud GPU instances) for model Deploy, scale, and monitor ML/LLM models in production using tools such as Triton Inference Server, vLLM, Optimize GPU utilization, cost, and throughput across multi-tenant workloads; manage CUDA/driver/toolkit Collaborate with data science/ML engineering teams on MLOps pipelines — model versioning, experiment 5–10 years of hands-on DevOps/SRE/Infrastructure engineering experience, including at least 2–3 years in a senior or lead capacity. Deep expertise in Git-based workflows and repository management at scale (GitHub/GitLab/Bitbucket). Proven experience designing observability stacks (Prometheus, Grafana, ELK/EFK, Datadog, New Relic, OpenTelemetry). Strong background in cloud platforms (AWS, Azure, and/or GCP) and on-premise/hybrid infrastructure. Expert-level skills with Infrastructure as Code (Terraform, Ansible, CloudFormation, or Pulumi). Strong Kubernetes and Docker experience, including multi-cluster and multi-environment management. Hands-on experience building and maintaining CI/CD pipelines end to end. Working knowledge of GPU infrastructure (NVIDIA CUDA, drivers, NCCL) and experience deploying ML/AI models to production. Proficiency in scripting/automation languages: Python, Bash, and/or Go. Solid understanding of networking, load balancing, DNS, and security fundamentals in distributed systems. Experience partnering with architects and engineering leads to translate designs into infrastructure and delivery plans. Hands-on with SAST, DAST, and SCA tooling (e.g., SonarQube, Snyk, Checkmarx, OWASP DependencyCheck) integrated directly into CI/CD pipelines. Container and image security: vulnerability scanning (Trivy, Grype, Clair), minimal/hardened base images, and signed/verified image provenance (Cosign/Sigstore). Familiarity with Secrets management and credential hygiene using tools such as HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault Excellent communication skills and comfort operating cross-functionally with development, data science, and product teams. Show more