GCP / SRE / DevOps Engineer
Cittabase Solutions
Role tags
Tech stack mentioned
Role overview
Formatting this description...
GCP / SRE / DevOps Engineer is a senior technical leadership role responsible for the reliability, scalability, and automation of a cloud-native GCP ecosystem. You will operate complex distributed systems built on GKE and SpringBoot microservices, drive toil elimination through advanced automation, and bridge application development with platform stability — ensuring a resilient, secure, and fully pipeline-driven environment. Key Responsibilities Leadership & Programme Management Act as the primary liaison between the client and Cittabase engineering teams, driving alignment on priorities, SLAs, and delivery commitments. Allocate and prioritize Jira bugs, incidents, and stories across support engineers; track team deliverables and enforce SLA adherence. Own and present weekly, monthly, and quarterly operational metrics reports to client stakeholders. Provide technical mentorship to the team; unblock engineers through hands-on debugging, architecture guidance, and escalation support. Incident Management & Site Reliability Serve as Incident Commander for high-priority outages — coordinate cross-functional response, lead blameless post-mortems, and drive systemic prevention measures. Drive incident management activities ensuring timely resolution; participate in on-call rotations via PagerDuty. Define, monitor, and enforce SLOs, SLIs, and error budgets across streaming and batch data platforms. Implement distributed tracing (OpenTelemetry) to diagnose latency and message loss; establish DORA metrics and Four Golden Signals (Latency, Traffic, Errors, Saturation) as operational baselines. Automate operational runbooks to reduce manual toil and mean time to recovery. Observability & Cost Optimisation Design and maintain Grafana and Cloud Monitoring dashboards tracking the Four Golden Signals and platform health across all workstreams. Monitor and optimise GCP cloud costs using billing insights and FinOps best practices. Platform Engineering & Infrastructure Design and manage production-grade GKE clusters ensuring high availability for SpringBoot microservices; implement and optimise cloud-native persistence using AlloyDB. Configure and maintain Apigee gateways for secure, low-latency API management. Manage the health and scaling of Kafka brokers and Pub/Sub topics to ensure zero message loss in streaming pipelines. Support operational health of large-scale data processing tooling: BigQuery, Dataflow, and Cloud Composer orchestration. Own end-to-end infrastructure lifecycle via Terraform — establish reusable module standards and state management for consistent environments across the GCP ecosystem. CI/CD & DevOps Automation Architect and maintain robust CI/CD pipelines using GitHub Actions; transition manual deployment processes into fully automated, gated workflows for Cloud Functions, Dataflow, and Cloud Composer. Qualifications 7+ years of experience in SRE or DevOps, with at least 4 years of hands-on leadership within the Google Cloud Platform (GCP). Proven experience in a leadership capacity during critical outages (Incident Commander) and a strong background in post-mortem documentation and remediation. Proven track record of building unified dashboards in Grafana and managing complex alerting rotations in PagerDuty. Expert-level knowledge of Kubernetes (GKE), including service mesh, ingress controllers, and cluster security. Advanced mastery of Terraform, GitHub Actions, and Kafka infrastructure management. Deep familiarity with GCP’s data and serverless offerings, including AlloyDB, Dataflow, Cloud Functions, and BigQuery. GCP Professional Cloud DevOps Engineer or Professional Cloud Architect certification is highly desirable. Show more