Posted today · be early
Site Reliability Engineer (SRE)
Novacard
WorldwideremotePosted today
Skill Required
Site-Reliability-EngineeringDevOpsInfrastructure-EngineeringCloud-EngineerSRESite-Reliability-EngineerSite-Reliability-Engineering-(SRE)Site-Reliability-Operations-EngineerDevOps-Site-Reliability-EngineerSite-Reliability-Engineering-JobsSite-Reliability-Engineer-IISite Reliability EngineerSystem DesignSOCObservabilityKubernetesPrometheusnetworkingautomationTerraformbuildingPythondesignNetwork MonitoringCI/CDDesign PatternsShell ScriptingAWSandElasticsearchGoCICDFulltime
Key highlights
- Fluent Russian required; English B1+ minimum.
- 3+ years of SRE, DevOps, or Infrastructure Engineering experience required.
- Fully remote work.
- Official employment under Russian Labor Code for Russia-based residents.
- Hands-on experience with Grafana and Zabbix required.
- Opportunity to work on a new digital product for the Mexican market.
Role overview
NOVACARD is the first interest-free and no-annual-fee credit card company in Mexico, designed to simplify personal finances and give users complete control via a mobile app — offering up to $200,000 MXN in credit with full digital management. We're looking for a Site Reliability Engineer (SRE) to ensure the stability, performance, and reliability of our critical production systems, working at the intersection of development and operations by building automation tools, improving observability, and preventing incidents before they occur.
Responsibilities
- Ensure the stability, performance, and fault tolerance of production systems.
- Develop and maintain infrastructure automation and observability tools.
- Monitor system health, respond to incidents, and perform root cause analysis (RCA).
- Collaborate with development teams to improve scalability and reliability of services.
- Define and manage SLIs, SLOs, and Error Budgets.
- Lead incident response: organize recovery, document RCA, and run blameless post-mortems.
- Configure and administer Grafana and Zabbix, design insightful dashboards, and fine-tune alerting.
- Integrate and monitor external vendor systems, collaborating with vendor technical support when needed.
Requirements
- Fluent Russian, English B1+ (comfortable with technical documentation).
- 3+ years of experience as an SRE, DevOps, or Infrastructure Engineer.
- Strong understanding of observability principles (metrics, logs, traces).
- Hands-on experience with Grafana and Zabbix (administration, configuration, alert optimization).
- Experience working with AWS and CI/CD tools.
- Practical knowledge of SLI/SLO/Error Budget frameworks.
- Experience leading and documenting incidents and post-mortems.
- Scripting skills for automation (Python, Bash, or Go).
- Solid understanding of distributed systems and networking fundamentals.
Nice to have
- Experience monitoring and supporting mobile applications.
- Familiarity with Terraform, Prometheus, Loki, ELK, or similar tools.
- Experience working with Kubernetes and containerized environments.
Benefits
- Fully remote work format.
- Official employment under the Russian Labor Code (for residents of Russia); contractor collaboration available for candidates from other countries.
- Opportunity to work in an international team on a new digital product for the Mexican market.
- A data-driven environment where your contributions have a real impact.
Additional details
- NOVACARD is the first interest-free and no-annual-fee credit card in Mexico.
- Users can access up to $200,000 MXN in credit, only pay when they use it, and manage everything digitally in under 5 minutes.
- Originally posted on Himalayas.