The DevOps Engineer – L1 is an entry-level, hands-on role responsible for independently owning infrastructure components, maintaining CI/CD pipelines, managing observability systems, and supporting operational reliability across production environments. This execution and ownership-focused position involves operating with minimal guidance, leading small-to-medium infrastructure improvements, and serving as the first point of escalation for junior engineers and interns. Over time, the role grows into driving platform initiatives, leading incident response, and contributing to architecture decisions.
Responsibilities
- Own and operate EKS cluster health, node group scaling, and namespace-level resource management across Dev, Staging, and Prod
- Manage and troubleshoot AWS services including EC2, RDS, IAM, VPC, S3, ALB, and CloudWatch, CloudTrail
- Write, review, and apply Terraform changes for infrastructure provisioning and modification
- Maintain and extend Jenkins pipelines and shared libraries; enforce security and quality gates across CI/CD stages
- Manage image promotion across QA → Staging → Prod using Git and GitHub as the source of truth
- Build and maintain Grafana dashboards using Prometheus metrics, Loki logs, and Mimir for long-term storage
- Own L1/L2 on-call rotation via PagerDuty; lead incident triage, drive resolution, and write postmortems
- Tune alerting rules and recording rules to reduce noise and maintain reliable signal
- Investigate security events, triage CVEs, and drive remediation through the vulnerability management workflow
- Maintain runbooks and identify opportunities to automate repetitive operational tasks
- Mentor DevOps interns on tooling, debugging, and operational best practices
Requirements
- Education: B.Tech or Dual Degree (Computer Science/IT preferred)
- Experience: 2-4 Years of relevant experience
- Hands-on experience with AWS core services (EC2, EKS, IAM, VPC, S3, ALB)
- Kubernetes: deployments, services, namespaces, RBAC, HPA, and basic troubleshooting
- Jenkins: building and maintaining pipelines and shared libraries
- Terraform: writing, reviewing, and applying infrastructure changes
- Git and GitHub: branching strategies, PR workflows, branch protection rules
- Grafana and Prometheus: building dashboards and writing PromQL queries
- Loki: log querying and aggregation (LogQL basics)
- PagerDuty: on-call participation, escalation policies, and incident management
- Linux administration and Bash scripting
- Python scripting for automation (boto3, requests, API integrations)
- Incident triage: log correlation, root-cause analysis, and postmortem writing
- Networking fundamentals: DNS, TCP/IP, HTTP/S, load balancers, and firewalls
Nice to have
- Mimir: long-term metrics storage and query optimisation
- Istio service mesh: AuthorizationPolicy, mTLS, Telemetry API
- FinOps basics: cost anomaly detection, tagging, rightsizing
Additional details
- Location: Bangalore
- Tools: Version control: Git, GitHub; CI/CD: Jenkins; Cloud: AWS; Infrastructure as Code: Terraform; Container orchestration: Kubernetes (EKS); Observability: Grafana, Prometheus, Loki, Mimir; Incident management: PagerDuty
- Persona traits: Operationally Confident, Systems Thinker, Automation-Minded, Security-Aware, Process-Oriented, Collaborative, Eager to Grow