Skip to main content
Zobhira
Home
Jobs
Certifications
Zobhira
JobsCertificationsTodayAbout
Log inSign up
Zobhira

New job and contest openings, updated every morning on one searchable board.

Find work

  • All jobs
  • Fresher roles
  • Remote roles
  • Certifications

Compete

  • Added today

Popular cities

  • India
  • Bangalore, Karnataka, India
  • Hyderabad, Telangana, India
  • Pune, Maharashtra, India
  • Mumbai, Maharashtra, India
  • Chennai, Tamil Nadu, India

Company

  • About
  • Contact
  • Privacy
  • Terms

Stay updated

One email a week with new roles.

Secure infrastructure
Free to use, no account needed to search

Board updated daily · © 2026 Zobhira. All rights reserved.

Privacy PolicyTerms of Service
Home / Jobs / Nexcess
Posted today · be early

Reliability Operations Specialist

Nexcess

IndiaremotePosted today
N

Skill Required

Reliability-Operations-SpecialistSite-Reliability-EngineerOperations-SpecialistIncident-ManagerIT-Operations-SpecialistReliability-SpecialistReliability-EngineerMaintenance-Reliability-EngineerProduction-Reliability-EngineerAsset-Reliability-EngineerSystems-Reliability-EngineerSite-Reliability-Operations-EngineerOperationsSystem DesignSOCObservabilityEngineeringnetworkinganalyticalautomationsecuritywrittendesignDevOpsLinuxCloudITILJiraandFulltime

Key highlights

  • $85,000 - $100,000 Annually
  • 3+ years of experience
  • Remote
  • Traditional and Roth 401(k) with company matching
  • Proficiency with Jira Service Management (JSM)

Role overview

Nexcess is looking for a Reliability Operations Specialist to drive operational excellence across incident management, service reliability, observability, and continuous improvement initiatives. This role serves as a central coordinator and subject matter expert for reliability practices, partnering with engineering, infrastructure, security, and operations teams to improve service stability, reduce operational risk, and strengthen operational readiness. This role does not have direct people management responsibilities but influences reliability outcomes through collaboration, process ownership, and data-driven decision making. It involves overseeing incident response, post-incident reviews, and providing visibility into platform health.

Responsibilities

  • Participate in major incident response activities and serve as an Incident Commander when assigned.
  • Facilitate incident coordination, escalation, stakeholder communications, and status reporting during service-impacting events.
  • Support ongoing improvement of incident management processes, procedures, and operational readiness.
  • Drive initiatives focused on reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Maintain and apply the Criticality Matrix to tier services, infrastructure, and customer MRR impact.
  • Coordinate and facilitate post-mortem reviews following significant incidents.
  • Ensure post-mortems are completed accurately, consistently, and within established timelines.
  • Synthesize findings across incidents to identify trends, recurring issues, and systemic risks.
  • Maintain accountability for corrective action tracking and closure.
  • Promote a blameless culture of learning and continuous improvement and proactive/reactive problem management.
  • Partner with engineering teams to define and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) across our product lines and services.
  • Support development and evolution of platform observability strategies, including monitoring, alerting, dashboards, and telemetry standards.
  • Analyze reliability metrics and operational trends to identify improvement opportunities.
  • Recommend and track initiatives that improve platform stability, resiliency, and service performance
  • Partner with DevOps to design JSM workflows for Incident, Problem, and Change processes while eliminating manual meetings through automation.
  • Establish change policy, lifecycle rules, risk assessments, CAB oversight, and approvals across standard, normal, and emergency changes.
  • Track and govern Change Failure Rates and change-related incidents to protect environment stability.
  • Embed across Product Line pods to manage service acceptance, operational readiness, support models, and lifecycle status (supported/unsupported/EOL).
  • Maintain asset governance, CMDB accuracy, Configuration Items (CIs), and service relationship mapping in JSM
  • Develop reliability reporting for engineering leadership and executive stakeholders.
  • Maintain incident communication standards and stakeholder notification protocols.
  • Provide regular reporting on reliability trends, corrective actions, incident performance, and service health.
  • Translate technical reliability metrics into actionable business insights.

Requirements

  • 3+ years of experience in Product Operations, Platform Operations, Technical Customer Support, Incident Coordination, Site Operations, IT Service Management (ITSM), or a related operational role.
  • Experience participating in or coordinating major incident response activities.
  • Knowledge of incident management, root cause analysis, problem management, and post-mortem methodologies.
  • Experience working with monitoring, alerting, observability, or operational reporting tools.
  • Strong analytical and organizational skills with exceptional attention to detail.
  • Excellent written and verbal communication skills.
  • Ability to work effectively across multiple teams and influence outcomes without direct authority.
  • Strong problem-solving skills and the ability to remain calm and organized during high-pressure situations
  • Hands-on proficiency with Jira Service Management (JSM) workflows, automation, and data quality governance.

Nice to have

  • Experience working with Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability metrics.
  • Familiarity with Linux systems, cloud infrastructure, networking concepts, hosting platforms, or distributed systems.
  • Knowledge of ITIL, operational excellence frameworks, or Site Reliability Engineering (SRE) principles.
  • Experience supporting high-availability SaaS, hosting, cloud, or infrastructure environments.
  • Experience creating executive-level operational reports, dashboards, and presentations.
  • Experience using observability and incident management platforms.

Benefits

  • Comprehensive benefits package
  • Traditional and Roth 401(k) with company matching
  • A collaborative, team-oriented culture
  • Consistent and predictable work hours
  • Engaging, varied work that keeps each day different
  • Opportunities to contribute ideas and influence how work gets done

Additional details

  • Location: Remote
  • Employment type: Permanent, Full-time
  • Pay Range: $85,000 - $100,000 Annually, The final compensation offered will be determined based on factors including location, experience, skills, qualifications, and market conditions.
  • Disclaimer: This job description is only a summary of the typical functions of the position. It is not intended to be an exhaustive or comprehensive list of all job responsibilities, tasks, or duties. Additional duties and tasks may be assigned as part of the job function. Nexcess reserves the right to modify, interpret, or apply this job description in a way that best supports the organizational needs. The job description in no way creates or implies an employment contract. The employment contract remains “at will”.
  • Equal Employment Opportunity Policy: Nexcess is committed to offering equal employment opportunity without regard to age, color, disability, gender, gender identity, genetic information, marital status, military status, national origin, race, religion, sexual orientation, veteran status, or any other legally protected characteristic.
  • Originally posted on Himalayas

Similar jobs open now

InfoBeans logo
QA Architect
InfoBeans
Indore, Madhya Pradesh, India
View details
Luxoft logo
Senior Data Engineer (immediate joiner)
Luxoft
India
View details
Confidential logo
Lead DevOps Engineer
Confidential
Ahmedabad, Gujarat, India
View details
Teamware Solutions logo
GCP Data Engineer – Python/BigQuery/Airflow/Terraform
Teamware Solutions
Bangalore, Karnataka, India
View details
Apply now
LocationIndia
TypeFulltime
SalaryUSD 85000 - 100000
Posted9/27/2026
Apply by11/26/2026

Links are checked every day. If this one stops working, tell us and we'll pull it.

More like this

View all
InfoBeans logo
QA Architect
InfoBeans
Luxoft logo
Senior Data Engineer (immediate joiner)
Luxoft
Confidential logo
Lead DevOps Engineer
Confidential
Apply now