DevOps Engineer | Engineering Team Manager , Vice President
BlackRock
Mumbai, IndiaPosted 1 month ago
Skill Required
DevOps EngineerEngineering ManagerEngineeringDevOpsBackup and RecoverytroubleshootingObservabilityautomationfintechdesignandAIFulltime
Key highlights
- 10+ years of experience required.
- Bachelor’s degree in computer science/engineering or equivalent.
- Flexible Time Off (FTO) offered.
- Hybrid work model: minimum 4 days in office, 1 day remote.
- Strong retirement plan, tuition reimbursement, and comprehensive healthcare benefits.
- Role focuses on AI‑enabled, autonomous infrastructure and service operations.
Role overview
You can work with us at one of top FinTech companies. We sell our Aladdin platform to over 200 of the top global corporations, in total managing about quarter of all the world’s money under management. BlackRock is global but close‑knit team of individuals who share a common goal of providing the very best possible level of support to our business partners and customers. From the top of the firm down, we embrace the diversity of values, identities and ideas brought by our employees. We are serious about our people and offer Flexible Time Off, collaborative working spaces and several other benefits.
Responsibilities
- Cover business‑critical computing workloads, real‑time / interactive processing, data transfer services, application and new technology on‑boarding and upgrades, and recovery procedures.
- Provide 24*7*365 support as part of an international team split into 4 global regions.
- Accountable for enterprise‑scale reliability, risk posture, and operational resilience through the adoption of AI‑enabled, autonomous infrastructure and service operations.
- Transition the organization from human‑driven, reactive operations to AI‑supervised, self‑healing systems, ensuring technology platforms scale safely, predictably, and in alignment with business growth and regulatory expectations.
- Own strategy, outcomes, and governance.
- Own reliability outcomes – deliver measurable improvements in availability, latency, recovery time, and incident recurrence for critical workloads.
- Establish and mature SRE practices: SLIs/SLOs, error budgets, and reliability governance.
- Build AI‑enabled, signal‑driven operations – replace alert volume with AI‑correlated signals that prioritize true business impact and early risk indicators.
- Implement and improve detection, correlation, and routing workflows integrated into operational processes and tooling.
- Engineer self‑healing systems (automation‑first) – design, implement, and govern automated remediation for known failure patterns (with safe guardrails and audit trails).
- Maintain structured human oversight for novel scenarios and ensure continuous learning feeds back into automation.
- Embed reliability and operability into engineering lifecycle – partner with Engineering/Architecture/Product to build operability, observability, resilience‑by‑design into services.
- Drive root‑cause elimination and reduce recurrence through systemic fixes, not repetitive recovery.
- Own change‑aware resilience + risk posture – use AI risk signals to anticipate failures based on historical data, dependency graphs, and weak points.
- Support production readiness: capacity planning, disaster recovery exercises, and disciplined change governance.
- Ensure AI‑driven systems are observable, explainable, and auditable, meeting operational and regulatory expectations.
- Develop and lead high‑performing global team, fostering strong ownership, technical depth, and a culture of accountability and continuous improvement.
- Collaborate with skilled professionals across the globe and manage a broad range of technologies and applications to deliver service quality and excellence.
Requirements
- Bachelor’s degree in computer science /engineering (or equivalent practical experience).
- 10+ years across Service Management, DevOps, SRE, Product Engineering, and/or large‑scale production operations.
- Strong hands‑on experience with observability/monitoring/telemetry platforms, focused on actionable insights and reliability outcomes.
- Proven experience transitioning environments from reactive support to proactive, signal‑driven / AI‑assisted operations.
- Designed/tuned/governed automation and AIOps workflows, enabling automated remediation while retaining structured human oversight for exceptions.
- Experience implementing change‑aware operations, drift detection/correction, and data‑driven reliability governance to reduce incident recurrence.
- AI‑assisted capacity forecasting and proactive scaling for performance predictability and cost efficiency.
- End‑to‑end operational fluency: telemetry → ITSM integration → automated execution.
- Experience sponsoring or governing AI‑assisted/autonomous operational platforms.
Benefits
- Flexible Time Off (FTO).
- Strong retirement plan.
- Tuition reimbursement.
- Comprehensive healthcare.
- Support for working parents.
- Collaborative working spaces.
- Hybrid work model – at least 4 days in the office per week, with the flexibility to work from home 1 day a week.
Additional details
- The Service Management Operations Group is responsible for monitoring, supporting, and administering production environments for all BlackRock businesses (including subsidiaries and BlackRock Solutions) acting as a first responder relative to troubleshooting, problem resolution, and escalation.
- The Operations Group delivers service quality and excellence through teamwork, innovating operational processes, and being part of the One BlackRock culture.
- Guidance on AI use for candidates – BlackRock encourages thoughtful AI usage for learning and preparation but focuses interview evaluation on candidates’ own experiences, thinking, and judgment.
- At BlackRock, the mission is to help more people experience financial well‑being; clients save for retirement, education, homes, and businesses, strengthening the global economy.
- BlackRock is proud to be an Equal Opportunity Employer, evaluating qualified applicants without regard to age, disability, family status, gender identity, race, religion, sex, sexual orientation and other protected attributes.
- The hybrid work model is designed to enable a culture of collaboration and apprenticeship while supporting flexibility for all employees.