Posted today · be early
Kubernetes Engineer
Kubegrade
IndiaremotePosted 1 day ago
Skill Required
Kubernetes-EngineerPlatform-EngineeringSite-Reliability-EngineeringSRE-DevOpsCloud-Infrastructure-EngineeringAzure-Kubernetes-EngineerKubernetes-Operations-EngineerKubernetes-Solutions-EngineerKubernetes-Platform-EngineerKubernetes-EngineeringKubernetes-Security-EngineerOperationsKubernetestroubleshootingObservabilityPrometheusnetworkingautomationsecuritydesignDevOpsArgoCDCloudHelmandAIFulltime
Key highlights
- Remote (global) position
- Full-time
- AI agents and GitOps workflows for Kubernetes automation
- Mandatory pre-interview technical assignment on their platform
- Strong Plus experience: operating 10+ clusters
- Strong Plus experience: Kubernetes upgrades across multiple versions
Role overview
Kubegrade is a company that automates Kubernetes upgrades, maintenance, drift detection, and troubleshooting using AI agents and GitOps workflows. They build for platform teams running real production clusters at scale. This is a full-time, remote position for a Kubernetes Engineer who will work hands-on with their platform to improve cluster visibility, troubleshooting, and operational workflows. The role assumes deep, practical Kubernetes experience and involves significant independent work validating production workflows.
Responsibilities
- Design, operate, and troubleshoot Kubernetes clusters in production environments
- Work directly with Kubegrades platform to validate real-world Kubernetes workflows
- Improve cluster visibility and UX
- Define troubleshooting and remediation logic
- Translate Kubernetes operational pain into product requirements
- Handle upgrades and API deprecations
- Address configuration drift
- Resolve misconfigured workloads, networking, storage, or security
- Contribute to internal tooling, automation, and best practices
- Act as a Kubernetes domain expert in product discussions
Requirements
- Strong, practical Kubernetes experience (not theoretical)
- Debug production incidents
- Reason about control plane vs workload issues
- Understand real upgrade and lifecycle challenges
- Deep knowledge of Kubernetes internals
- Experience with Helm, Kustomize, or GitOps workflows
- Experience with observability stacks (Prometheus, Grafana, logs, traces)
- Comfortable working autonomously in a remote, async environment
Nice to have
- Platform / SRE / DevOps background in regulated or large-scale environments
- Experience operating 10+ clusters
- Kubernetes upgrades across multiple versions
- API deprecation management
- Multi-cloud or hybrid environments
- Opinionated views on how Kubernetes tooling should work
Additional details
- Access and create account on app.kubegrade.com
- Connect at least one Kubernetes cluster (cloud, on-prem, or local) before any interview
- Explore cluster and namespace visualization
- Review signals, issues, and context surfaced by the platform
- Test AI agents for troubleshooting and remediation actions
- Document what the product does well
- Identify what is unclear or broken
- Assess how it fits (or doesnt) into real Kubernetes operations
- Propose how you would improve it for platform teams
- Pre-interview assignment serves as a mandatory filter
- The interview is a technical discussion centered on: Your Kubernetes experience, Your hands-on usage of Kubegrade, Your ability to critique, reason, and propose improvements